Instructions to use HuggingFaceBio/Carbon-A-1.2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use HuggingFaceBio/Carbon-A-1.2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="HuggingFaceBio/Carbon-A-1.2B", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForTokenClassification model = AutoModelForTokenClassification.from_pretrained("HuggingFaceBio/Carbon-A-1.2B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Carbon-A-1.2B
A DNA annotation model from the Carbon family.
Carbon-A predicts protein-coding sequence (CDS) at single-base resolution in eukaryotic genomes. It produces separate forward- and reverse-strand probabilities, preserving coding regions that overlap on opposite strands. The annotation pipeline in this repository turns these probabilities into gene models with a confidence score per gene.
Facts
- 1.2B-parameter DNA annotation model with two strand-specific classification heads.
- Tokenizer: non-overlapping 6-mers; each DNA token represents six bases.
- Sequence length: 16,384 tokens (98,304 bp). The annotation pipeline reads longer records in windows that overlap by 1,024 tokens (6,144 bp).
- Inputs: genomic DNA. The annotation pipeline reads FASTA and GenBank files, optionally gzip-compressed.
- Outputs: per-base probabilities for non-coding and CDS on each strand. The annotation pipeline turns them into gene models with a confidence score per gene.
- Training data: 2,055 RefSeq GCF assemblies (2,047 species) from mammals, other vertebrates, invertebrates, plants, fungi, and protozoa.
- Transformers interface:
AutoModelForTokenClassificationwithtrust_remote_code=True; Transformers 4.56 or later, including 5.x. FlashAttention-2 needs Transformers 4.x. - Inference: BF16 with FlashAttention-2, the setting of all reported results. PyTorch SDPA (
--attn-implementation sdpa) needs no FlashAttention installation, but its probabilities are not bit-identical.
How to use
Install
The annotation pipeline runs on Linux. It needs Python 3.10 or later, a CUDA build of PyTorch, and a C++ compiler with OpenMP. If you do not have a suitable compiler, we recommend GCC, for example on Debian or Ubuntu:
sudo apt install build-essential
Then download this repository and install its requirements:
pip install -U huggingface_hub
hf download HuggingFaceBio/Carbon-A-1.2B --exclude model.safetensors --local-dir Carbon-A-1.2B
pip install -r Carbon-A-1.2B/requirements.txt
pip install --no-build-isolation flash-attn
The weights (4.6 GB) are downloaded to the Hugging Face cache on first use.
Annotate a genome
bash Carbon-A-1.2B/annotation_pipeline/run_annotation.sh --input genome.fna.gz
The input is a FASTA or GenBank file, optionally gzip-compressed. Carbon-A reads it in 98,304-bp windows that overlap by 6,144 bp and averages the overlapping predictions. The pipeline then decodes the gene structures and writes the results to results/<accession>/:
prediction.gff: genes, mRNAs, and CDS. Every gene carriesconfidence, the probability that its structure is exactly right.prediction.fnaandprediction.faa: the CDS and protein sequences.*.raw_probs.parquet: the per-base CDS probabilities.
All genes are written. The reported results count the genes with a confidence of at least 0.1.
The standard genetic code is used by default. For another code, pass --codon-table. Tetrahymena thermophila (GCF_000189635.1) uses translation table 6 automatically. A likely codon-table mismatch is detected automatically: the pipeline adds a warning to results/WARNINGS.txt and renames the result directory to results/<accession>_WARNING.
The pipeline README lists all options. It also describes the evaluation against a reference annotation and BUSCO.
Test on a benchmark genome
The benchmark genomes and their NCBI annotations are in the data bucket under eval/. To test the pipeline on yeast, download its genome and annotation:
url=https://huggingface.co/buckets/HuggingFaceBio/Carbon-A-training-data/resolve/eval/seen/GCF_000146045.2
mkdir -p GCF_000146045.2
for file in GCF_000146045.2_R64_genomic.fna genomic.gff cds_from_genomic.fna sequence_report.jsonl; do
curl -fsSL -o GCF_000146045.2/$file $url/$file
done
Annotate the genome and compare the result with the NCBI annotation:
bash Carbon-A-1.2B/annotation_pipeline/run_annotation.sh \
--input GCF_000146045.2/GCF_000146045.2_R64_genomic.fna
bash Carbon-A-1.2B/annotation_pipeline/run_evaluation.sh \
--prediction-prefix results/GCF_000146045.2/prediction \
--genome-dir GCF_000146045.2
results/GCF_000146045.2/evaluation_metrics.json then reports a nucleotide F1 of 0.998, an exon F1 of 0.943, and a gene F1 of 0.953.
License
Apache 2.0.
- Downloads last month
- 29