graph LR
SkbioObject["SkbioObject"]
Sequence["Sequence"]
GrammaredSequence["GrammaredSequence"]
MetadataMixin["MetadataMixin"]
IntervalMetadata["IntervalMetadata"]
DissimilarityMatrix["DissimilarityMatrix"]
DistanceMatrix["DistanceMatrix"]
TreeNode["TreeNode"]
GeneticCode["GeneticCode"]
skbio_util["skbio.util"]
SkbioObject -- "provides" --> Sequence
SkbioObject -- "provides" --> DissimilarityMatrix
SkbioObject -- "provides" --> TreeNode
Sequence -- "extends" --> SkbioObject
GrammaredSequence -- "extends" --> Sequence
MetadataMixin -- "enhances" --> Sequence
IntervalMetadata -- "supports" --> Sequence
DistanceMatrix -- "extends" --> DissimilarityMatrix
DissimilarityMatrix -- "extends" --> SkbioObject
TreeNode -- "extends" --> SkbioObject
GeneticCode -- "supports" --> GrammaredSequence
skbio_util -- "supports" --> SkbioObject
skbio_util -- "supports" --> Sequence
skbio_util -- "supports" --> MetadataMixin
skbio_util -- "supports" --> DissimilarityMatrix
skbio_util -- "supports" --> TreeNode
The Core Data Structures & Utilities component in scikit-bio serves as the foundational layer, providing the essential building blocks for representing biological data and offering common utilities that ensure consistency and reusability across the library. This component is critical for maintaining a cohesive and robust architecture, as it defines the fundamental types and behaviors upon which more complex bioinformatics functionalities are built.
The most fundamental base class in scikit-bio. It provides a common interface and basic functionalities (like equality testing and hashing) for many scikit-bio objects, ensuring consistency across diverse data types.
Related Classes/Methods:
SkbioObject(1:1)
The abstract base class for all biological sequences (e.g., DNA, RNA, Protein). It defines core sequence-agnostic operations such as slicing, concatenation, and basic metadata handling.
Related Classes/Methods:
Sequence(1:1)
Extends Sequence by incorporating a defined grammar (alphabet, degenerate characters, gap characters). This class is crucial for handling sequences with specific biological alphabets and rules, like DNA or Protein sequences.
Related Classes/Methods:
GrammaredSequence(1:1)
A mixin class that provides generic metadata handling capabilities. Objects inheriting from this mixin can store and manage arbitrary key-value pair metadata, allowing for rich annotation of biological data.
Related Classes/Methods:
MetadataMixin(1:1)
A specialized data structure for managing metadata associated with specific intervals (e.g., genomic regions) on a sequence or other linear data. It allows for precise annotation of sub-regions.
Related Classes/Methods:
IntervalMetadata(1:1)
The base class for representing pairwise dissimilarities (distances) between samples. It provides methods for accessing, manipulating, and validating dissimilarity data, which is central to many ecological and phylogenetic analyses.
Related Classes/Methods:
DissimilarityMatrix(1:1)
A subclass of DissimilarityMatrix specifically designed for symmetric distance data, where the distance from A to B is the same as B to A. It adds specific validation for symmetry.
Related Classes/Methods:
DistanceMatrix(1:1)
The core data structure representing a node within a phylogenetic tree. It supports tree construction, traversal, and manipulation, forming the basis for all phylogenetic analyses.
Related Classes/Methods:
TreeNode(1:1)
Represents a genetic code, mapping codons to amino acids. This utility is essential for translating nucleotide sequences into protein sequences.
Related Classes/Methods:
GeneticCode(1:1)
A module containing a collection of general-purpose helper functions, decorators, and testing utilities used across the library. This includes functions for validation, common data manipulations, and internal helpers.
Related Classes/Methods:
skbio.util(1:1)