PROTECT YOUR DNA WITH QUANTUM TECHNOLOGY
Orgo-Life the new way to the future Advertising by AdpathwayAn Artificial-Intelligence System Can Turn Viral Sequences into Evolutionary Maps
When a new virus begins spreading, one of the most urgent scientific questions is not simply what the pathogen is, but how it is changing. Researchers need to determine whether newly collected genetic sequences belong to a known lineage, whether they represent a distinct strain, and how they relate to viruses detected elsewhere. That process traditionally requires several separate stages of bioinformatics, specialist judgment and manual editing. A study published in Neural Computing and Applications now describes an end-to-end artificial-intelligence system designed to automate much of that workflow, from reading a nucleotide sequence to producing a simplified, publication-quality phylogenetic tree. The researchers tested the approach using Orthoavulavirus 1, the virus commonly known as Newcastle disease virus, or NDV, as a proof of concept.
NDV is an important test case because it affects domestic and wild birds and can cause serious outbreaks in poultry. Its evolutionary history is complex, with viral sequences collected across different countries and time periods. Classification is often based on genetic features of the fusion, or F, gene, which encodes a protein involved in the virus’s ability to enter host cells. Specific changes in that gene can help distinguish viral groups and are central to the established phylogenetic classification of NDV. By focusing on the F gene, the system targets a biologically meaningful region rather than treating the genome as an undifferentiated string of letters. The researchers say the same general design could eventually be adapted to other viruses and genomic markers.
The pipeline begins with a nucleotide sequence supplied by a user. That sequence may represent a complete viral genome or only a partial genome, a practical feature because samples obtained during surveillance are not always fully sequenced. The program first searches for open reading frames, or ORFs—the stretches of nucleotides that can potentially be translated into proteins. In molecular biology, identifying ORFs is a foundational step in sequence annotation because genes are encoded as instructions that begin with start signals, continue through a reading frame and terminate at stop codons. Locating these regions allows the software to identify candidate viral genes, including the NDV F gene, before downstream classification and evolutionary analysis take place.
The F-gene sequence is then passed to an AI model trained to classify NDV sequences. Although the study’s abstract does not report a single accuracy figure, the model forms the central recognition component of the system, distinguishing sequence patterns associated with viral categories. The authors describe an approach built with modern machine-learning tools, drawing on techniques associated with deep learning and sequence analysis. In practical terms, such a model can learn statistical relationships among nucleotide positions from a reference dataset rather than relying only on manually written rules. That does not mean the model “understands” viral evolution in a biological sense. Instead, it identifies recurring patterns in examples whose classifications are already known, and uses those patterns to assign a classification to an incoming sequence.
Classification alone, however, cannot show how viruses are related to one another. To reconstruct those relationships, the system automatically builds a phylogenetic tree using the maximum-likelihood method. Phylogenetics treats DNA or RNA sequences as records of evolutionary history: mutations accumulate over time, and closely related viruses are expected to share more genetic changes than distant ones. Maximum likelihood evaluates alternative tree structures and models of sequence evolution, selecting the arrangement that makes the observed genetic data most probable under the chosen assumptions. Before that calculation, sequences generally must be aligned so that comparable nucleotide positions are placed in the same columns. The study’s reference list identifies MAFFT for multiple-sequence alignment and established maximum-likelihood software and methods as part of the computational foundation.
A raw phylogenetic tree can be scientifically informative but difficult to interpret, particularly when it contains large numbers of viral genomes. Each tip may represent an individual isolate, while internal branches indicate inferred relationships and branch lengths can reflect the amount of genetic change. When hundreds or thousands of sequences are included, labels overlap, long branches dominate the display and the larger evolutionary structure becomes hard to see. The new system addresses this problem with what the researchers call multi-level simplification. Rather than requiring an analyst to prune and reorganize the tree manually, the software progressively condenses the result into more readable forms while retaining key relationships. It also generates accompanying metadata tables, helping users connect branches to information such as collection location and sampling date.
The ability to filter a tree by geography or time could be especially useful during an outbreak. A global tree may reveal broad evolutionary groupings, but public-health investigators often need a narrower view: which viruses were detected most recently, whether sequences from a particular region cluster together, or how a local sample fits into an international pattern. The supplementary material includes an interactive version of a simplified tree based on the 15 latest viruses, along with a corresponding metadata table. These features could help transform a dense genomic dataset into a visual summary that is easier for epidemiologists, veterinary authorities and other decision-makers to inspect. The system is not presented as a replacement for laboratory confirmation or expert interpretation, but as a way to reduce the delay between sequence generation and an actionable evolutionary analysis.
The researchers assembled the application as a linked workflow with a user interface and backend components, aiming to make the process accessible to people who may not routinely write code or operate specialist phylogenetic software. The underlying implementation draws on Python-based scientific and bioinformatics libraries, sequence-processing tools and machine-learning frameworks. The source code is available through a public GitHub repository, while the study’s supplementary files provide a model dataset containing GenBank accession numbers used for training, interactive tree visualizations and metadata outputs. Public code and data are important for assessing systems of this kind because reproducibility depends on more than a polished interface. Independent researchers need to examine the training data, test the model on sequences collected outside its original dataset and determine how sensitive the pipeline is to incomplete, low-quality or unusual genomes.
That scrutiny will be essential before an automated classifier can be relied upon in fast-moving outbreaks. Machine-learning models may perform well on sequences resembling their training data but behave less predictably when confronted with a novel lineage, sequencing errors or a virus that has recombined or evolved in an unexpected way. Phylogenetic trees also depend on choices about alignment, evolutionary models, sampling and data quality; a visually attractive tree can still convey unwarranted confidence if those assumptions are not examined. The NDV demonstration therefore represents a proof of concept rather than evidence that the platform is ready for every pathogen or surveillance setting. Even so, automating the repetitive steps of ORF extraction, target-gene classification, tree construction and visualization could give specialists more time to focus on biological interpretation and outbreak response.
As viral sequencing becomes faster and more widespread, the bottleneck in surveillance is increasingly the conversion of raw genetic data into understandable evidence. The system reported by Ansar Yousif and colleagues aims to narrow that gap by combining neural-network classification with conventional evolutionary inference and automated visualization. Its most distinctive contribution is not simply the use of AI to label a sequence, but the attempt to connect that prediction to a complete analytical chain that ends with a simplified tree and contextual metadata. If validated across larger datasets and additional viral species, tools built on this model could help laboratories and public-health agencies trace viral movement and evolution with less manual intervention. For now, the Newcastle disease virus application illustrates both the promise and the limits of algorithmic genomics: machines can rapidly organize biological information, but scientists remain responsible for testing the result, understanding its uncertainty and deciding what it means in the real world.
Subject of Research: Automated artificial-intelligence classification of viral nucleotide sequences and generation and simplification of phylogenetic trees, demonstrated with Orthoavulavirus 1 (Newcastle disease virus)
Article Title: Automated virus classification and phylogenetic tree generation
Article References: Yousif, A., Abdrabo, S., Elsayed, R. et al. “Automated virus classification and phylogenetic tree generation.” Neural Computing and Applications 38, 702 (2026). Original research article
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12362-y
Keywords: viral classification, phylogenetic analysis, deep learning, Newcastle disease virus, Orthoavulavirus 1, viral evolution, sequence annotation, outbreak surveillance
Tags: AI system for viral phylogeneticsAI-based virus evolution analysisautomated viral genome analysisbioinformatics workflow automationevolutionary mapping of infectious virusesgenetic features of viral fusion geneNewcastle disease virus genetic sequencingphylogenetic tree constructionviral lineage identificationviral mutation and strain differentiationviral outbreak tracking using artificial intelligenceviral sequence classification


3 hours ago
5




















English (US) ·
French (CA) ·