Artificial intelligence and bioproduction: when data becomes a performance driver
Bioproduction is undergoing a major transformation. As biological processes become increasingly complex, automation and artificial intelligence are opening up new possibilities: better understanding cell behaviour, anticipating deviations and, ultimately, making processes more predictive. But this evolution depends on one essential factor: generating data that is genuinely relevant.
From vaccines and monoclonal antibodies to insulin, cell therapies and gene therapies, bioproduction uses living organisms — bacteria, yeasts or cells — to manufacture products of interest. And its applications extend far beyond pharmaceuticals, reaching cosmetics, food production, flavours and proteins.
Complex biological processes to control
Producing with living organisms involves multiple stages: selecting cell lines, developing the bioprocess, choosing culture media, adjusting temperature, agitation and oxygenation, and finally scaling up to industrial production.
Moving from a bioreactor of just a few millilitres to a tank holding several thousand litres is not simply a matter of multiplying volumes. Geometry, agitation systems, pumps, heat transfer and compound diffusion all change with scale. Each step can therefore affect cell behaviour and process productivity.
This variability can have significant consequences. A failure at laboratory scale may mean lost time and samples; at industrial scale, it can mean losing an entire batch, with major financial consequences and potentially an impact on patient supply.
Lots of data — but is it really useful?
Bioreactors are already highly instrumented. Temperature, pH, dissolved oxygen, agitation, glucose and lactate can all be monitored throughout a culture. The issue is therefore not necessarily a lack of data, but its biological value.
Measuring temperature or pH every second generates huge amounts of information, but does not necessarily explain why cells stop multiplying, enter apoptosis or produce less of the target biomolecule.
To make artificial intelligence truly useful, laboratories therefore need to go beyond accumulating physicochemical data and identify indicators that provide direct information about cell status and process evolution. This is closely linked to the Quality by Design approach: understanding critical parameters in order to better control production.
Looking for information inside the cell
One promising approach is to monitor cellular biomarkers. Analysing messenger RNA, for example, can provide information about oxidative stress, proliferation, energy production and responses to infection or transfection, as well as the precursors of the product being manufactured.
This changes the logic of process monitoring. Rather than simply observing that a process has deviated, it becomes possible to investigate the biological mechanisms behind that deviation.
However, this requires extremely fast and reproducible sampling. Some biomarkers change very rapidly, meaning that a manually collected sample analysed several hours later may no longer accurately represent the state of the cells at the time of collection.
Automating sampling to improve data quality
Automation therefore becomes an essential part of the process. Robotic systems can perform regular sampling 24/7 while reducing both sample volumes and human intervention.
Another objective is to limit the shear stress imposed on cells during handling. Traditional pumps and sampling methods can alter cell condition and introduce bias into subsequent analyses. Fluidic and acoustic technologies can instead be used to capture, concentrate and prepare cells while reducing this impact.
Most importantly, automation makes it possible to generate longitudinal datasets: following the same bioproduction process over time and observing precisely how it evolves.
Deep learning enters laboratory equipment
Artificial intelligence can also be integrated directly into laboratory instruments. Deep learning image analysis can, for example, automatically identify and count cells despite variations in size and morphology.
In the system presented during the session, the algorithm can adapt sampling according to the required number of cells. Once the target quantity has been detected, sampling stops automatically.
AI is therefore no longer used only to analyse experimental results afterwards: it becomes part of the instrument’s operation itself.
From troubleshooting to digital twins
The next step is to combine these different sources of information: bioreactor data, physicochemical parameters, omics data, biomarkers and production conditions.
Initially, these datasets can help laboratories understand what happened. Why did a culture fail? Was it oxidative stress, an energy issue, poor induction or another process deviation?
With enough repeated experiments, the same data can then be used to predict.
This is where the concept of the digital twin comes into play: creating a virtual model of a bioprocess capable of simulating the effect of a change before applying it to real production.
Changes in agitation speed, temperature or injection methods could therefore be evaluated digitally to anticipate their impact on cells and production yield.
Useful AI starts with relevant data
Artificial intelligence is not a magic solution to the complexity of bioproduction. Its performance depends directly on the quality, diversity and relevance of the data used to train and operate it.
The challenge for laboratories is now to move from simply monitoring a process to genuinely understanding the living systems within that process.
Automation, biomarkers, image analysis, omics data and predictive models could help reduce failures, accelerate development and improve scale-up. Ultimately, the objective is clear: more reliable, more efficient and better-controlled bioproduction.
Article based on the session dedicated to artificial intelligence applied to bioproduction, organised by Polepharma at Forum LABO.

