How is AI transforming biotechnology?

From sequence to bioprocess: why artificial intelligence and Europe’s research infrastructures will shape the next decade of biomanufacturing
Marco Anteghini1 (BIOS project) and Vitor A. P. Martins dos Santos1 (IBISBA)
1 Lifeglimmer GmbH, Berlin
Introduction
During the past decade, artificial intelligence (AI) has increasingly become a routine companion for biotechnologists. It accelerates bioprocess development (the design and optimisation of biological production methods), supports strain and process engineering (e.g. modifying organisms and processes to improve efficiency), and streamlines fermentation monitoring (e.g. tracking biological reactions that produce desired molecules). AI facilitates the integration and analysis of large, complex datasets, thus providing improved biological insights (Holzinger et al., 2023; Anteghini et al., 2025; Asín-García et al., 2025).
Biotechnology is inherently interdisciplinary. It translates life science knowledge into industrial and societal applications. Progress in the field is increasingly driven by advances in modelling (e.g. creating digital representations of biological systems), automation, data science, and bioengineering (applying engineering principles to biological systems).
IBISBA, the Infrastructure for Biotechnology Services for Biomanufacturing, and partner initiatives such as the BIOS project investigate the impact of AI, automation and data-driven approaches on research and innovation in industrial biotechnology and biomanufacturing. A concrete example is the BIOS bio-intelligent Design-Build-Test-Learn (DBTL) cycle. Here, AI, biosensors (devices that detect biological changes), automation and digital twins (virtual representations of physical systems) combine to accelerate strain and process engineering in Pseudomonas putida producer strains (types of bacteria engineered for production). These strains are used to produce terpenes (aromatic compounds), polyolefins (a type of plastic), and methyl acrylate (a valuable chemical building block). In this setting, AI is not only used to analyse data, but also to support hybrid learning across cellular- and process-level models. This helps researchers move faster from strain design to bioprocess optimisation (improving the production process).
To implement AI and digital twins in laboratory settings, teams typically assemble high-quality, standardised datasets; select and integrate sensors for real-time data collection; develop or adopt machine-learning (ML) models tailored to specific bioprocesses; establish digital infrastructure for data management and model deployment; and connect AI-driven recommendations to automated systems or decision workflows. Researchers must understand both the opportunities and challenges of this transition to shape European research strategy and support a competitive, sustainable circular bioeconomy (European Commission, 2024). This article provides an overview of evidence for AI-driven biotechnology and the supporting infrastructure required.
A growing role in the literature
The increasing role of AI in biotechnology is evident in the scientific literature. A PubMed search for publications combining AI or ML with biotechnology shows a marked increase, from approximately 240 articles in 2013 to over 3,200 in 2025, representing an eightfold rise since 2015 alone (Figure 1).

This increase reflects a substantive change in practice. ML and deep learning (DL) now support tasks previously limited by speed and scalability. These tasks include process data integration, experimental design, strain engineering, pathway modelling, fermentation monitoring, and extraction of information from sensors and online analytics (Anteghini et al., 2025; Asín-García et al., 2025; Müller et al., 2026). In both laboratory and industrial settings, AI accelerates data analysis, enhances process understanding, and streamlines research and development pipelines.
The practical wins are concrete. ML-guided strain and enzyme engineering uses AI to design or modify microorganisms and proteins, thereby reducing the experimental search space prior to large-scale screening campaigns (Vanella et al., 2022). Soft sensors, which are software-based tools, can estimate process variables (such as temperature, pH, or chemical concentration) in real time. These variables are otherwise difficult, slow, or expensive to measure directly (Hikosaka et al., 2020). Hybrid models integrate mechanistic knowledge (understanding of the biological or chemical processes at work) with data-driven prediction. They support process optimisation, scale-up and control (Narayanan et al., 2022; Faure et al., 2023). Neural-mechanistic approaches use neural networks in combination with mechanistic models, showing how genome-scale metabolic models (comprehensive representations of an organism’s metabolic processes) can be combined with ML to improve predictive power while retaining biological interpretability (Faure et al., 2023). Together, these approaches reduce the time from a biological concept to a validated biomanufacturing process.
More recently, generative AI (genAI) has begun to extend this transformation beyond data analysis. GenAI can help orchestrate complex computational workflows. For example, it translates bioprocess questions (issues related to the biological manufacturing process) into executable pipelines (automated sequences of computational tasks), connects experimental metadata (data describing experiment conditions) to modelling tools, generates surrogate models (simplified mathematical models that approximate complex processes), and systematically explores design hypotheses (possible designs to test in experiments) (Huntington et al., 2023). In engineering biology, this strengthens the Design-Build-Test-Learn cycle (an iterative engineering process for improving biological systems) by reducing technical barriers. It also enables broader engagement with advanced modelling, simulation and digital process design (Gurdo et al., 2023).
Interdisciplinary collaborations, digital twins and the questions they raise
The integration of AI and biotechnology promotes interdisciplinary collaboration. This spans computational sciences, life sciences, process engineering, automation, ethics and industrial strategy. One important outcome is the digital twin. In this context, it is a virtual model of a bioprocess that updates with experimental or real-time data. Researchers can use it to test operating strategies, media compositions, feeding regimes and scale-up options in silico. Digital twins can reduce the number of costly pilot runs and support more informed process decisions (Asín-García et al., 2025; Müller et al., 2026).
When connected across sites, digital twins also enable distributed learning. It is a process in which computer models are trained locally at each site without sending all sensitive data to a central server. These locally trained models can then be compared, updated, or combined (aggregated) to improve performance across different types of equipment, locations, raw materials (feedstocks), and regulatory environments (San et al., 2023). This is not only a technical shift; it also creates a new collaboration model, where knowledge can move across a network of participants while data ownership, confidentiality, and intellectual property are managed under clear rules (governance).
This approach is particularly relevant for decentralised and distributed biomanufacturing. In a circular bioeconomy, producers may shift production from a small number of centralised plants to flexible networks. These networks adapt to local biomass availability, regional waste valorisation opportunities, demand fluctuations, and supply-chain disruptions. AI-supported digital twins, federated data systems and modular bioprocess infrastructures facilitate monitoring, comparison, optimisation, and transfer of processes across sites.
Achieving this vision requires generating data that machines can read, harmonise, standardise, and align with FAIR principles: findable, accessible, interoperable, and reusable (Wilkinson et al., 2016). This requirement compels infrastructures such as IBISBA to invest in metadata catalogues, data spaces, and knowledge hubs. These frameworks facilitate the exchange of data, models, and protocols under defined governance. They enable sharing while safeguarding sensitive information. As a pan-European distributed research infrastructure, IBISBA lays the foundation for reusable, interoperable, and AI-ready biomanufacturing services (IBISBA, n.d.).
The same transformation raises important concerns. Digitalisation and automation introduce uncertainties around data governance, interoperability, security, ethics, resilience, and human oversight. Müller et al. (2026) caution that automation erodes tacit knowledge and safety oversight. Even tools that support the green transition may carry hidden sustainability or equity risks. Regulatory and design frameworks such as Safe and Sustainable by Design must adapt to this bio-digital convergence (Reins & Wijns, 2025). Shared FAIR infrastructures will prevent siloed, redundant, and poorly governed data systems.
An outlook built on shared European infrastructure
As biotechnology evolves, initiatives such as IBISBA seek to integrate expertise and equipment across Europe into coordinated, interoperable research infrastructure. IBISBA envisions digitalised innovation ecosystems in a distributed, highly interoperable network. This network can function as a digital twin factory. It consists of reusable components for data access, metadata capture, modelling, simulation, monitoring and validation. This is distinct from isolated, single-purpose models.
IBISBA currently provides a unified access point to integrated services for end-to-end bioprocess development. The IBISBA Knowledge Hub supports the organisation, management and sharing of research data. Bioindustry 4.0 and related initiatives are also developing services for real-time online monitoring, high-quality datasets, digital twins, online sensors and data-driven methods in industrial biotechnology (Bioindustry 4.0, n.d.). These components enable the practical implementation of smart biomanufacturing.
If these foundations are built, entirely new operating models become plausible. Flexible production networks, on-demand manufacturing and decentralised biomanufacturing systems could be digitally coordinated. They can adapt to local feedstocks, regional infrastructure and changing demand. The objective is to achieve scalable, resilient, and affordable biomanufacturing that responds more effectively to societal needs.
Future efforts will need to strengthen collaboration among research communities, industry and policy, while also developing the interdisciplinary skills required for digital biomanufacturing. Researchers and technical staff will increasingly need expertise in data science, automation engineering, systems biology, bioinformatics, digital infrastructure management, regulatory frameworks and responsible technology deployment. Skills in managing and interpreting large datasets, designing digital workflows, operating laboratory automation systems, applying machine learning, and complying with FAIR data standards are becoming essential (OECD, 2020). Training programmes that connect life scientists, engineers, data specialists and infrastructure operators will help prepare the workforce for these new roles.
Conclusion
The integration of AI into biotechnology is not a marginal upgrade. It is a structural shift in how research and innovation are conducted, from how we design a strain to how we monitor, optimise and scale a bioprocess. The publication trend makes the direction unmistakable. The open questions make equally clear that technology alone is not enough.
Continued collaboration, shared standards, FAIR data practices and open knowledge exchange across Europe will be key to responsibly leveraging these opportunities. The forthcoming joint event aims to foster this kind of cross-project and cross-disciplinary dialogue among infrastructures, researchers and industry.
References
Anteghini, M., Gualdi, F., & Oliva, B. (2025). How did we get there? AI applications to biological networks and sequences. Computers in Biology and Medicine, 190, 110064. DOI: 10.1016/j.compbiomed.2025.110064.
Asín-García, E., Fawcett, J. D., Batianis, C., & Martins dos Santos, V. A. P. (2025). A snapshot of biomanufacturing and the need for enabling research infrastructure. Trends in Biotechnology, 43(5), 1000-1014. DOI: 10.1016/j.tibtech.2024.10.014.
Bioindustry 4.0. (n.d.). Bioindustry 4.0: Research infrastructures for smart biomanufacturing.
European Commission. (2024). Building the future with nature: Boosting biotechnology and biomanufacturing in the EU. COM(2024) 137 final.
European Commission, CORDIS. (n.d.). BIOS: The bio-intelligent DBTL cycle, a key enabler catalysing the industrial transformation towards sustainable biomanufacturing. Project ID 101070281.
Faure, L., Mollet, B., Liebermeister, W., & Faulon, J.-L. (2023). A neural-mechanistic hybrid approach improving the predictive power of genome-scale metabolic models. Nature Communications, 14, 4669. DOI: 10.1038/s41467-023-40380-0.
Gurdo, N., Volke, D. C., McCloskey, D., & Nikel, P. I. (2023). Automating the design-build-test-learn cycle towards next-generation bacterial cell factories. New Biotechnology, 74, 1-15. DOI: 10.1016/j.nbt.2023.01.002.
Hikosaka, T., Aoshima, S., Miyao, T., & Funatsu, K. (2020). Soft sensor modeling for identifying significant process variables with time delays. Industrial & Engineering Chemistry Research, 59(26), 12156-12163. DOI: 10.1021/acs.iecr.0c01655.
Holzinger, A., Keiblinger, K., Holub, P., Zatloukal, K., & Müller, H. (2023). AI for life: Trends in artificial intelligence for biotechnology. New Biotechnology, 74, 16-24. DOI: 10.1016/j.nbt.2023.02.001.
Huntington, T., Baral, N. R., Yang, M., Sundstrom, E. R., & Scown, C. D. (2023). Machine learning for surrogate process models of bioproduction pathways. Bioresource Technology, 370, 128528. DOI: 10.1016/j.biortech.2022.128528.
IBISBA. (n.d.). Infrastructure for Biotechnology Services for Biomanufacturing.
Martins dos Santos, V., Anton, M., Szomolay, B., et al. (2022). Systems Biology in ELIXIR: Modelling in the spotlight. F1000Research, 11, 1265. DOI: 10.12688/f1000research.126734.2.
Müller, A., Robaey, Z., Youssef, S., Martins dos Santos, V. A. P., & Asín-García, E. (2026). Digitalisation and automation in smart biomanufacturing: Uncertainties, challenges and early pathways towards safe and sustainable design. New Biotechnology, 93, 378-391. DOI: 10.1016/j.nbt.2026.04.007.
Narayanan, H., Luna, M. F., Sokolov, M., Butté, A., & Morbidelli, M. (2022). Hybrid models based on machine learning and an increasing degree of process knowledge: Application to cell culture processes. Industrial & Engineering Chemistry Research, 61(25), 8658-8672. DOI: 10.1021/acs.iecr.1c04507.
OECD. (2020). Building digital workforce capacity and skills for data-intensive science. OECD Science, Technology and Industry Policy Papers, No. 90. DOI: 10.1787/e08aa3bb-en.
Reins, L., & Wijns, J. (2025). The “Safe and Sustainable by Design” concept: A regulatory approach for a more sustainable circular economy in the European Union? European Journal of Risk Regulation. DOI: 10.1017/err.2024.25.
San, O., Pawar, S., & Rasheed, A. (2023). Decentralized digital twins of complex dynamical systems. Scientific Reports, 13, 20087. DOI: 10.1038/s41598-023-47078-9.
Vanella, R., Kovacevic, G., Doffini, V., Fernández de Santaella, J., & Nash, M. A. (2022). High-throughput screening, next generation sequencing and machine learning: Advanced methods in enzyme engineering. Chemical Communications, 58, 2455-2467. DOI: 10.1039/D1CC04635G.
Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J. J., et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. DOI: 10.1038/sdata.2016.18.
