Analytical chemists have been using machine learning long before ChatGPT made headlines. Now, as generative AI enters the laboratory, the discipline faces both opportunities and risks

  • AI is already deeply embedded in analytical chemistry, particularly through machine learning and chemometrics, where it helps scientists analyse large datasets, identify patterns and improve the speed and rigour of spectroscopic and other analytical techniques.
  • Generative AI is expanding the field’s capabilities, enabling researchers to create synthetic spectra, predict molecular structures and explore new chemical hypotheses, although experts stress that the technology is still developing and has not yet reached its full practical potential.
  • Researchers caution against overreliance on AI outputs, noting that both chatbots and scientific AI models can produce convincing but incorrect results or ‘hallucinations’, making validation against established physical and chemical principles essential.
  • AI-powered analytical tools are finding real-world applications in areas such as forensic science, cancer diagnostics and scientific translation, while experts generally believe AI will augment rather than replace analytical chemists by shifting their role towards interpretation, quality control and oversight.

This summary was generated by AI and checked by a human editor

If you were to ask people to name the branch of science or technology that they expect to impact their lives most in the near future, for good or ill, there’s a very good chance that they will choose artificial intelligence. 2024’s Nobel prize bonanza for AI pioneers – the chemistry prize for protein structure prediction with AlphaFold, and the physics one for the ‘foundational discoveries that led to artificial neural networks’ – occurred a bare two years after the chatbot ChatGPT burst onto the scene, threatening to upend higher education and much else besides. Like all professionals, chemists are rapidly coming to terms with living and working in what is fast becoming an AI world. This is perhaps particularly true for analytical chemists, whose work revolves around identifying and quantifying the composition of substances.

If you then ask someone what AI means to them, many would describe ChatGPT or a similar chatbot. This, however, is only one small part of the technology. AI is defined more generally as ‘the capacity of a computer to perform activities that are normally associated with human reasoning’ – learning, problem-solving, perception, decision-making and the like – and in that general sense it has been around almost as long as computers themselves.

And this is not even the first period that has seen explosive growth in the apparent potential of AI; rather, it is the third, and the first two didn’t last. Business analysts use a framework known as the Gartner hype cycle to describe how novel technologies can move from a ‘peak of inflated expectations’ through the inevitable ‘trough of disillusionment’ to a final ‘plateau of productivity’. The first peak in expectations for AI came in the 1950s with the first, very simple neural networks; Isaac Asimov’s short stories about humanoid robots, I, Robot, were published together in 1950. The second coincided with the growth of machine learning techniques in the 1980s, and the third, of course, with generative AI and chatbots today.

AI’s long road in chemistry

So, what will the current technology wave look like in its plateau of productivity? Rasmus Bro, who researches machine learning in analytical chemistry at the University of Copenhagen in Denmark, stresses that for him, the plateau is still some way off. ‘We routinely analyse samples using machine learning techniques that were introduced in the 1980s, and that is all we need,’ he explains. ‘Generative methodologies will not necessarily help us do what we do much better, but they will broaden the kind of problems that we can solve … we think it will be a revolutionary change, not an incremental one when it comes, but we’re not there yet,’

Long before generative AI emerged onto the scene, established forms of AI, particularly machine learning, had revolutionised how scientists approach data analysis. The power of machine learning to classify and discern patterns in large datasets makes analysis simultaneously more rigorous and less time-consuming. Any chemist who uses large quantities of data – which means almost any chemist today, including many students – will be using machine learning, whether they realise it or not. In particular, machine learning aids Bro’s discipline of chemometrics, defined as ‘the science of extracting information from chemical data using statistical and mathematical methods’, and reinforces the value of the spectroscopies and other analytical techniques that provide that data.

We should beware in particular of a ”beautifully correct” answer from AI

Generative AI is, essentially, a particular type of machine learning. The difference between it and other types lies in that word ‘generative’. Unlike traditional machine learning algorithms, generative algorithms can generate something entirely novel from analysing the patterns stored in their vast datasets. Any generative model is based on a probabilistic data model and can generate new examples based on that probability. Jerome Workman Jr., a former instrument and software development scientist from California, US, says generative models will ‘augment, simulate, and better characterise spectral data’. This provides the logical link between scientific uses of generative AI and the ubiquitous chatbots: in ChatGPT, for instance, text-based large language models – in some cases derived from ‘the whole [accessible] Internet’ – are equivalent to repositories of scientific data.

However, using machine learning to generate original content from a prompt or request is neither (quite) as new or as strange as has been suggested. Farooq Wahab, an analytical chemist and research engineering scientist at the University of Texas at Arlington in the US, has tracked down what may be its first mention, from over 30 years ago. ‘To the best of my knowledge, the first use of the term “generative AI” with something like its current meaning was in a talk by a British-American forensic software analyst, Andy Johnson Laird, in a conference paper about AI and intellectual property in 1991,’ he says.

Wahab takes a firmly critical attitude to the use of generative AI in his own field, explaining his approach using a ‘clever Hans’ analogy. Clever Hans was a horse that appeared to do simple sums during exhibitions in early 20th century Germany. Later, a psychologist, Oskar Pfungst, showed that Hans was, in fact, responding to involuntary cues from his owner. By analogy, artificial intelligence programs – including generative ones – will sometimes give a plausible answer through flawed logic. ‘We should beware in particular of a ”beautifully correct” answer from AI, because it may still be based on incorrect reasoning,’ he adds.

Spectroscopy and AI at a crossroads

In very general terms, a typical experiment in analytical spectroscopy involves passing a beam of radiation through a sample, recording the amount absorbed at each wavelength and generating a spectrum for analysis and interpretation. The mathematical techniques used for interpreting the spectra form an important part of the discipline of chemometrics. At the very simplest level, this might just involve calibrating an instrument by simple regression, but more complex levels of spectral interpretation require machine learning techniques for descriptive and predictive analysis.

A recent feature in Spectroscopy magazine described spectroscopy as ‘at a crossroads’. The authors list three unrelated trends as contributing to the challenges facing spectroscopists: artificial intelligence (not necessarily limited to generative AI) – is one, of course, along with automation and miniaturisation. The trend towards using spectroscopy tools outside large-scale laboratory facilities – in the clinic, perhaps, or in the field for forensic applications – is a key factor driving miniaturisation.

ChatGPT is still not as good at chemistry as it is at maths

Workman believes that generative AI became particularly compelling for spectroscopists when it was able to offer the possibility of mapping data space itself, not just mapping inputs (the spectroscopic data) to outputs (the predicted analyte concentrations or values). ‘These [AI-generated] maps or models become important whenever it is difficult, expensive, time-consuming or perhaps impossible to obtain representative samples for calibration,’ he adds. ‘Importantly, we are able to generate physically plausible synthetic spectra using these techniques.’

All these techniques still have limitations, however. As many lecturers know, students who use large language models to help with their assignments are often tripped up by incorrect or ‘hallucinatory’ references: references that even experts in a field will find entirely plausible at first glance but that just don’t exist. ‘AI-generated spectra, too, may be “hallucinatory”; that is, plausible but incorrect, particularly if the algorithms have been poorly trained or are used outside the domain they were trained on,’ adds Workman. ‘But rigorous validation and cross-checking with the physical and chemical principles involved can provide safeguards.’

And ‘chatbots’ – basic text-based generative AI tools – can also prove helpful in spectroscopic analysis. Wahab has used ChatGPT as a coding assistant and found that it started off as an unreliable one but significantly improved. ‘ChatGPT is still not as good at chemistry as it is at maths,’ he explains. ‘I wouldn’t recommend it to a chemist for help with chemistry, but it can be a very useful PhD-level mathematical assistant for chemists without maths backgrounds. Or you could train it [in chemistry] as Omar Yaghi is doing.’ Yaghi, who won a share of the 2025 Nobel prize in chemistry for his work developing flexible and stable metal–organic frameworks, is a ‘power user’ of generative AI.

Promise, pitfalls and hallucinations

Without strict ethical safeguards, training chatbots with large bodies of data will come at a price. The AI company Anthropic ‘skimmed most of the world’s knowledge’ from the Internet to train its widely used chatbot, Claude; an enormous number of in-copyright books were accessed without permission for this, including some of Workman’s spectroscopy textbooks. ‘Five of my books are included in the list of works for a class action lawsuit suing the company for improper use,’ he explains. ‘This is such a large suit that each member could only get a small sum, but it’s not the money but the principle that’s important here.’

Schematic showing  Overview of SpectraML, translating between Spectrum Space and Molecule Space

Source: © 2025 Kehan Guo et al

Maching learning can help predict spectra from structures, and vice versa

Xiangliang Zhang and Kehan Guo are computer scientists who work closely with NMR spectroscopists and other chemist colleagues at the University of Notre Dame in the US state of Indiana to formulate spectroscopy-related problems that can be studied using machine learning and generative AI. One direction is to predict spectra from molecular structures; another, the more challenging inverse problem, is to generate candidate molecular structures that are consistent with a given spectrum.

Generative models are useful here because they can rapidly explore a much larger space of possible molecular candidates than would be practical through manual reasoning alone, giving chemists and spectroscopists a broader set of hypotheses to evaluate. These candidates may be represented as Smiles strings and compared with chemical databases such as PubChem as one part of the validation process, although the absence of an exact database match does not by itself establish novelty, synthesisability or chemical value. The goal is not to replace chemical expertise, but to help chemists focus their effort on assessing which AI-generated candidates are chemically valid, spectroscopically plausible, and experimentally meaningful.

AI beyond the laboratory

Any technology that can separate and characterise complex mixtures of compounds may find uses in forensic science, diagnostics and food safety, for example. Combining the basic techniques with AI can add accuracy, reliability, speed and rigour to already well-established methodologies. In some ways, this technology mimics a well-studied biological system: olfaction, our sense of smell. Mammals have evolved, and in the case of dogs near-perfected, a system for distinguishing between substances in the atmosphere from their molecular profile. Volatile organic compounds (VOCs) passing into mammalian noses bind to specific olfactory receptors and generate signals that are recognised as odours. A human nose is theoretically able to recognise about 10,000 different ones, and dogs, with 40 times as many olfactory receptors, can distinguish millions.

Close up of an 'electronic nose' - small electronic circuitry on gold-coloured metal - held in someone's fingertips

Source: © Olov Planthaber/Linköping University

Puglisi’s electronic nose uses machine learning algorithms to classify different VOCs

Donatella Puglisi, a physicist based at Linkoping University in Sweden, and her group have developed an artificial olfactory system, called an ‘e-Nose’, in which volatiles bind to a sensor array, generating signals that machine learning algorithms can classify quickly, precisely and from small samples. The e-Nose comprises an array of 32 different gas sensors held consistently at four different temperatures. Electrical signals generated when gas containing VOCs passes over the array are transferred to an advanced machine learning system. ‘The artificial intelligence in the e-Nose is trained to distinguish one pattern of signals from another, rather as our brains are trained in early childhood to distinguish the smell of cheese from that of strawberry jam,’ explains Puglisi.

Scheme showing an overview of the classification pipeline from signal acquisition to final output, featuring steps from 1) new sample, 2) new e-nose measurements, 3) 32 voltage-time signals, 4) feature engineering, 5) trained ML model, 6) intermediate predictions, 7) majority voting algorithm, 8) final decision, either post- or ante-mortem

Source: © 2025 Ivan Shtepliuk et al

The e-Nose and its algorithm aims to determine whether a sample was ante- or post-mortem

The sensitivity and precision of the e-Nose should be of particular value during murder investigations. Detectives need to be able to distinguish between tissue from living and dead individuals and between human and animal remains, and to estimate the time of death. This system is not yet available for use in the field, but it should offer a significant improvement over two currently used methods. The most appropriate standard analytical technique is gas chromatography–mass spectrometry (GC–MS), but this equipment is too fragile to use at a scene of crime, and the analyses are time-consuming; specially trained dogs have had major successes, but their results cannot be used as physical evidence in court.

Detailed spatiotemporal response of sensor #10 in the e-nose to VOCs emitted from blood plasma samples,  There are clear differences in the graphs between healthy individuals and cancer patients

Source: © 2026 Ivan Shtepliuk et al

Even an untrained eye can tell the difference between the results for healthy people (left) and cancer patients (right)

Beyond forensic science, the e-Nose is being tested in oncology, to distinguish between blood plasma from ovarian cancer patients and healthy controls. Puglisi is collaborating with her colleague, Jens Eriksson, and a company, VOC Diagnostics, to produce a version for clinical use. ‘We aim to produce a machine that can screen for ovarian cancer in minutes using a simple blood sample, with more accuracy than any other technique,’ she says. The company focuses on ovarian cancer primarily because its founder, György Horvath, is a gynaecologist who had specialised in this disease for decades, but the same principles could apply to any cancer type. ‘We would like to build a multi-cancer sensor, but first we need to make sure that different types of cancer “smell” differently,’ she adds.

Will AI replace analytical chemists?

Not every application of AI in general, and of generative AI in particular, that benefits chemists is specific to chemistry. It has already proved very useful for translating the scientific literature. English has only been the official language of science for a few decades, and only 40 years ago many chemistry students were required to know some German. Much valuable knowledge is hidden in papers written in other languages, and human translation is time-consuming and expensive. Google Translate and other standard translation software is useful but limited, particularly by an inability to handle mathematical symbolism. ‘Generative AI can do an excellent job of translating all types of chemistry papers … if the equations are provided in LaTeX format,’ says Wahab. This type of AI translation is making information in in the old multivolume Beilstein or Gmelin Handbooks of Organic Chemistry, or German quantum chemistry papers, for example, accessible to further generations of chemists.

But while AI in all forms is coming to play an ever-increasing role in the work of analytical chemists, one big question remains: will the discipline cease to be AI-enabled, and become AI-dominated? In fact, will human analytical chemists still be needed? Perhaps despite his involvement in the class action, Workman is an optimist on the jobs question. He suggests that the answer will be ‘more, not less’. ‘The more routine work will be automated,’ he explains. ‘But the analytical chemist’s role will switch to quality control and interpretation; demand will grow for scientists who understand chemometrics and can work with AI.’ If he is right, it seems that a bright future awaits at least some analytical chemists – and their robot collaborators.

Clare Sansom is a science writer based in Cambridge, UK