PlantQTLdb Assistant
DNA Analysis with Evo2
score supplied DNA, compare matched substitutions or request a model-generated continuation
Explicit Evo2 instruction + DNA; Research recommended
Open PlantQTLdb AssistantChoose the operation
These requests use the official Evo Designer website through the existing Assistant conversation. Explicit Evo2 requests containing valid DNA have their own sequence-analysis route; unlike the six new tools, they are not restricted to Research mode. Research is a convenient choice when working through this guide.
Paste DNA/FASTA or load a supported text sequence file into the message, then send it with an explicit Evo2 instruction. Availability depends on the official service.
How to ask and what you receive
Score DNA
Use Evo2 to score the following DNA sequence:
>example
ATGGGCGGTTCATGATGCGGCTGAAGCTGCCAAACGGTGTGACGACGAGCGAGCAGACGAGGTACCTGGCGAGCGTGATCGAGGCGTACGGCAAGGAGGGCTGCGCCGACGTGACAACCCGCCAGAACTGGCAGATCCGCGGCGTCACGCTCCCCGACGTGCC
what you receive
On successful completion, a model-labelled score summary with negative log-likelihood (NLL), per-base model tracks and available downloads. No score is promised here: the values must come from the actual service response.
Compare a substitution
Use Evo2 to compare the reference and alternative DNA sequences:
>reference
ATGGGCGGTTCATGATGCGGCTGAAGCTGCCAAACGGTGTGACGACGAGCGAGCAGACGAGGTACCTGGCGAGCGTGATCGAGGCGTACGGCAAGGAGGGCTGCGCCGACGTGACAACCCGCCAGAACTGGCAGATCCGCGGCGTCACGCTCCCCGACGTGCC
>alternative
ATGGGCGGTTCATGATGCGGCTGAAGCTGCCAAACGGTGTAACGACGAGCGAGCAGACGAGGTACCTGGCGAGCGTGATCGAGGCGTACGGCAAGGAGGGCTGCGCCGACGTGACAACCCGCCAGAACTGGCAGATCCGCGGCGTCACGCTCCCCGACGTGCC
what you receive
Scores for two matched sequences and their difference. Reference must be first, alternative second, with equal length and matching context at the first base. This example introduces one substitution; indels are not supported by this comparison.
Generate a continuation
Use Evo2 to generate a continuation of the following DNA with the Arabidopsis thaliana species condition:
>prompt
ATGGGCGGTTCATGATGCGGCTGAAGCTGCCAAACGGTGTGACGACGAGCGAGCAGACGA
what you receive
A model-generated 200-base continuation under the Arabidopsis condition if the website completes the task. This output is an unvalidated candidate, not an experimentally verified sequence or recovered reference annotation.
What the scores and plots mean
- NLL
- Lower mean negative log-likelihood means the observed bases are more expected in the model's context. There is no universal threshold for a good or functional sequence.
- Entropy
- Uncertainty in the predictive distribution; it is not a direct measurement of evolutionary conservation.
- Sequence logo
- Model probabilities normalized over A/C/G/T. Letter heights show model concentration, not experimentally measured motif enrichment.
- Comparison
- Alternative-minus-reference score differences depend on the model, orientation and context. They are not calibrated probabilities of pathogenicity or functional effect.
Input limits
Scoring accepts up to five DNA records, at most 16,000 bases each and 32,000 bases in total, subject to the input byte limit. The first base provides context and is not scored. Positions refer to the supplied sequence.
Generation requires one 2–500-base prompt and an explicit Arabidopsis thaliana or Zea mays condition. The generation condition does not turn sequence scoring into a species-specific annotation model.