Researchers at the University of Hong Kong (HKU) have developed a new AI algorithm that significantly improves the detection of cancer mutations using long-read DNA sequencing. The method, ClairS, uses synthetic tumour data to train a deep-learning model, enabling accurate analyses even when real patient data is limited.
According to the researchers, this approach could accelerate the application of precision medicine and improve the detection of genetic abnormalities. The results have been published in *Nature Methods*. ClairS is also available as open-source software via GitHub.
AI trained using synthetic tumour data
Accurately identifying mutations that occur exclusively in tumour cells is essential for cancer research and personalised treatments. However, many existing analysis methods have been developed for short-read sequencing, in which DNA is sequenced in short fragments. As a result, genetic abnormalities in complex parts of the genome sometimes go undetected. ClairS has been specifically developed for long-read sequencing, a technique that allows much longer DNA fragments to be analysed. This enables even hard-to-access regions of the genome to be mapped more effectively.
A key innovation of the system is the way in which the AI model is trained. As high-quality datasets containing tumour material are in short supply, the researchers developed a method that combines sequencing data from healthy human samples to create realistic, synthetic tumour-normal datasets. This creates a virtually unlimited amount of training data that simulates various scenarios, such as variations in tumour purity, sequencing depth and mutation frequency. According to the researchers, this results in a robust AI model that is better able to cope with the variation encountered in clinical practice.
High accuracy
ClairS’s performance was evaluated using datasets from various types of cancer, including breast cancer, lung cancer, melanoma and pancreatic cancer. In these cell models, the algorithm was able to identify small somatic mutations with a high degree of accuracy.
Lead researcher Ruibang Luo argues that long-read sequencing is radically transforming research into cancer genomes, particularly as genetic regions that were previously difficult to analyse are now becoming accessible. He believes that by using synthetic training data, a powerful AI model can be developed without relying on large quantities of real tumour data. The researchers view this approach as a scalable solution for the development of medical AI applications in situations where high-quality clinical datasets are scarce.
Encorporated in official workflow
The technology has now been incorporated into Oxford Nanopore Technologies’ official workflow for the analysis of somatic mutations. As a result, ClairS is already part of a commercial analysis tool for long-read sequencing, a significant step towards wider application within clinical genomics.
According to the researchers, the study demonstrates that synthetic data can play a key role in training AI models for medical applications. Not only could this lead to more accurate detection of cancer mutations, but the method also provides a foundation for future applications of long-read sequencing within oncological diagnostics and precision medicine. Further validation in clinical settings will be needed to determine the extent to which the technology can actually support the diagnosis and treatment of cancer patients.
European ambitions
During a debate organised by the European SYNTHIA project at the ICT&health World Conference in January, the role of synthetic data in AI-driven healthcare took centre stage. The participants concluded that artificially generated datasets could offer an important solution to the limited availability of health data in Europe, where privacy legislation and fragmented data sources hamper research. According to the European Commission, synthetic data can help with training AI models, testing digital healthcare solutions and accelerating drug development, without compromising patient privacy.
At the same time, researchers emphasised that synthetic datasets are only valuable if they are scientifically valid, clinically useful and resistant to being traced back to individual patients. They can also contribute to research into rare conditions and the creation of synthetic control groups in clinical trials. However, experts warned of risks such as the carry-over of existing bias from real-world datasets and uncertainty regarding liability when AI systems make errors. Clear validation methods, transparency and clear European guidelines are therefore seen as essential to increasing trust in synthetic data and its wider application within the healthcare sector.