danjwarr/gene-go-annotator-1b-single-term
Llama-3.2-1B-Instruct – QLoRA Fine-tuned for single gene identifier and ontology description prediction
This model is a QLoRA fine-tuned version of meta-llama/Llama-3.2-1B-Instruct.
Model description
This model adapts Llama-3.2-1B-Instruct using parameter-efficient fine-tuning (QLoRA) to extract gene identifiers and gene ontology descriptions from research articles. Specifcally for articles relating to organisms present in the Eukaryotic Pathogen, Vector and Host Informatics Resource (VEuPathDB).
Intended uses & limitations
Intended Uses
- Extracting gene identifier and gene ontology terms from research articles
- Intended for research and experimentation only
Limitations
- Trained on a task-specific dataset
- Not suitable for clinical, medical, or production use
- May produce incorrect or hallucinated outputs
- Inherits all limitations of the base Llama model
Training and evaluation data
Trained on the results section of publicly available PubMed research articles that were used for annotation of (VEuPathDB) genes. This version was trained under the supervised task of predicting single gene identifiers followed by gene ontolog terms based on the results sections of articles (i.e. given prompt & article text the model will produce geneId: go term descriptions).
Training procedure
Fine-Tuning Method
- QLoRA (4-bit quantization)
- Causal language modeling objective
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 5e-05
- trainbatchsize: 1
- evalbatchsize: 1
- seed: 42
- optimizer: Use OptimizerNames.ADAMWTORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizerargs=No additional optimizer arguments
- lrschedulertype: linear
- num_epochs: 3.0
Evaluations results
- Loss: 2.5472
- Accuracy: 0.5830
Framework versions
- Transformers 4.51.3
- Pytorch 2.5.1
- Datasets 3.5.0
- Tokenizers 0.21.1
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# Define parameters
max_input_size = 10240
temperature = 0.9
top_k = 50
top_p = 0.95
# Load model
tokenizer = AutoTokenizer.from_pretrained("danjwarr/gene-go-annotator-1b-single-term")
model = AutoModelForCausalLM.from_pretrained("danjwarr/gene-go-annotator-1b-single-term", torch_dtype=torch.float16)
# Prepare prompt
# Note: The model does not require this prompt at inference time
# The article text alone is sufficient; this is kept to match training
system_prompt = """<|begin_of_text|><|start_header_id|>system<|end_header_id|>You are an assistant tasked with extracting the gene function and location from biological literature only about [species]. You should ignore anything which describes a different species. You should describe any Biological Process, Molecular Function or Cellular Component related to the Gene product that has been identified. These descriptions should be one or two full sentences, in the same detail as given in the Gene Ontology. Some examples of entries in the Gene Ontology are <examples><example>. \"Entry of a symbiont into the body, tissues, or cells of a host organism as part of the symbiont life cycle. The host is defined as the larger of the organisms involved in a symbiotic interaction.\" </example> <example>. \"The secretion of neuropeptides contained within a dense core vesicle by fusion of the granule with the presynaptic membrane, stimulated by a rise in cytosolic calcium ion concentration </example> <example>. A heterotetrameric protein complex that associates with replication origins, where it is required for the initiation of DNA replication, and with replication forks </example> <example> Combining with the neurotransmitter dopamine and activating adenylate cyclase via coupling to Gi/Go to initiate a change in cell activity.</example> <example> Any process that modulates the frequency, rate or extent of the growth of all or part of an organism so that it occurs at its proper speed, either globally or in a specific part of the organisms development. </example> </examples><|eot_id|>\n<|start_header_id|>user<|end_header_id|>\nIn the following article, look for evidence of any Biological Processes, molecular functions, or cellular components that involve [species] and the gene product identified by [geneID]. Give a quote from the article of the supporting evidence as complete sentences. Supply these as a numbered list. If there is no evidence of any specific biological process, molecular function, or cellular component, reply with an empty list []. Add no other commentary. The article follows: """
# ***NEED INPUT HERE***
article_text = "" # replace this with research article text
# Create the full prompt
prompt = system_prompt + article_text + " <TERMS>" # system_prompt may be omitted
# Tokenize input and make a prediction
inputs = tokenizer(prompt, return_tensors="pt", max_length=max_input_size, truncation=True)
model_output = model.generate(
inputs["input_ids"],
max_new_tokens=100,
do_sample=True,
temperature=temperature,
top_k=top_k,
top_p=top_p,
num_return_sequences=1,
pad_token_id=tokenizer.eos_token_id,
attention_mask=inputs["attention_mask"]
)[0]
# Detokenize result
output_response = tokenizer.decode(model_output, skip_special_tokens=False)
if "<TERMS>" in output_response:
generated_output = output_response.split("<TERMS>")[1].split("<|begin_of_text|>")[0].strip()
# Print prediction
print(generated_output)