Academic Digital Library for institutions, students, and solo learners

Discovery

Content search and filters

This is the first search surface for the seven content types. Next we will connect full-text search and metadata-specific filters.

Results 1,892

Periodicals

From Plate to Production: Artificial Intelligence in Modern Consumer-Driven Food Systems

Global food systems confront the urgent challenge of supplying sustainable, nutritious diets in the face of escalating demands. The advent of Artificial Intelligence (AI) is bringing in a personal choice revolution, wherein AI-driven individual decisions transform food systems from dinner tables, to the farms, and back to our plates. In this context, AI algorithms refine personal dietary choices, subsequently shaping agricultural outputs, and promoting an optimized feedback loop from consumption to cultivation. Initially, we delve into AI tools and techniques spanning the food supply chain, and subsequently assess how AI subfields$\unicode{x2013}$encompassing machine learning, computer vision, and speech recognition$\unicode{x2013}$are harnessed within the AI-enabled Food System (AIFS) framework, which increasingly leverages Internet of Things, multimodal sensors and real-time data exchange. We spotlight the AIFS framework, emphasizing its fusion of AI with technologies such as digitalization, big data analytics, biotechnology, and IoT extensively used in modern food systems in every component. This paradigm shifts the conventional "farm to fork" narrative to a cyclical "consumer-driven farm to fork" model for better achieving sustainable, nutritious diets. This paper explores AI's promise and the intrinsic challenges it poses within the food domain. By championing stringent AI governance, uniform data architectures, and cross-disciplinary partnerships, we argue that AI, when synergized with consumer-centric strategies, holds the potential to steer food systems toward a sustainable trajectory. We furnish a comprehensive survey for the state-of-the-art in diverse facets of food systems, subsequently pinpointing gaps and advocating for the judicious and efficacious deployment of emergent AI methodologies.

Biotechnology2023arXiv
Periodicals

Novel transfer learning schemes based on Siamese networks and synthetic data

Transfer learning schemes based on deep networks which have been trained on huge image corpora offer state-of-the-art technologies in computer vision. Here, supervised and semi-supervised approaches constitute efficient technologies which work well with comparably small data sets. Yet, such applications are currently restricted to application domains where suitable deepnetwork models are readily available. In this contribution, we address an important application area in the domain of biotechnology, the automatic analysis of CHO-K1 suspension growth in microfluidic single-cell cultivation, where data characteristics are very dissimilar to existing domains and trained deep networks cannot easily be adapted by classical transfer learning. We propose a novel transfer learning scheme which expands a recently introduced Twin-VAE architecture, which is trained on realistic and synthetic data, and we modify its specialized training procedure to the transfer learning domain. In the specific domain, often only few to no labels exist and annotations are costly. We investigate a novel transfer learning strategy, which incorporates a simultaneous retraining on natural and synthetic data using an invariant shared representation as well as suitable target variables, while it learns to handle unseen data from a different microscopy tech nology. We show the superiority of the variation of our Twin-VAE architecture over the state-of-the-art transfer learning methodology in image processing as well as classical image processing technologies, which persists, even with strongly shortened training times and leads to satisfactory results in this domain. The source code is available at https://github.com/dstallmann/transfer_learning_twinvae, works cross-platform, is open-source and free (MIT licensed) software. We make the data sets available at https://pub.uni-bielefeld.de/record/2960030.

Biotechnology2022arXiv
Periodicals

Chemical communication between synthetic and natural cells: a possible experimental design

The bottom-up construction of synthetic cells is one of the most intriguing and interesting research arenas in synthetic biology. Synthetic cells are built by encapsulating biomolecules inside lipid vesicles (liposomes), allowing the synthesis of one or more functional proteins. Thanks to the in situ synthesized proteins, synthetic cells become able to perform several biomolecular functions, which can be exploited for a large variety of applications. This paves the way to several advanced uses of synthetic cells in basic science and biotechnology, thanks to their versatility, modularity, biocompatibility, and programmability. In the previous WIVACE (2012) we presented the state-of-the-art of semi-synthetic minimal cell (SSMC) technology and introduced, for the first time, the idea of chemical communication between synthetic cells and natural cells. The development of a proper synthetic communication protocol should be seen as a tool for the nascent field of bio/chemical-based Information and Communication Technologies (bio-chem-ICTs) and ultimately aimed at building soft-wet-micro-robots. In this contribution (WIVACE, 2013) we present a blueprint for realizing this project, and show some preliminary experimental results. We firstly discuss how our research goal (based on the natural capabilities of biological systems to manipulate chemical signals) finds a proper place in the current scientific and technological contexts. Then, we shortly comment on the experimental approaches from the viewpoints of (i) synthetic cell construction, and (ii) bioengineering of microorganisms, providing up-to-date results from our laboratory. Finally, we shortly discuss how autopoiesis can be used as a theoretical framework for defining synthetic minimal life, minimal cognition, and as bridge between synthetic biology and artificial intelligence.

Biotechnology2013arXiv
Periodicals

SOWAHA as a Cancer Suppressor Gene Influence Metabolic Reprogramming

SOWAHA is a protein-coding gene, also known as ANKRD43. Studies have indicated that SOWAHA can serve as a prognostic biomarker in colorectal cancer and pancreatic cancer. However, there are few reports about SOWAHA in other types of cancer and the specific mechanism of action of SOWAHA in cancer is also not clear. Based on National Center for Biotechnology Information (NCBI), The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression Project (GTEx), cBioPortal, Human Protein Atlas (HPA), etc., we adopted bioinformatics methods to uncover the potential tumor genomic features of SOWAHA, including the correlation with prognosis, gene mutation, immune cell infiltration, and DNA methylation in different tumors and evaluated the association with tumor heterogeneity, stemness, chemokines chemokine receptors, and immunomodulators in pan-cancer. Besides, we knocked down SOWAHA in SW620 cells and performed RNA-seq analysis, then we conducted functional enrichment to uncover the biological significance of the gene set. SOWAHA has early diagnostic potential, and low expression of SOWAHA was associated with poor prognosis in was associated with poor prognosis in GBMLGG, PAAD, READ, etc. SOWAHA is associated with most tumor immune-infiltrating cells in pan-cancer. SOWAHA correlates with DNA methylation, tumor heterogeneity, and stemness in many epithelial carcinomas. Furthermore, SOWAHA is involved in many enzyme activity and metabolic pathways, mainly metabolic programming pathways in cancer. Additionally, we identified two potential transcription factors of SOWAHA, TBX4, and FOXP2, which are dysregulated in SW620 cells. Besides, the cell proliferation and viability in siSOWAHA groups are better than in siNC groups.SOWAHA, identified as a suppressor gene, and its role in the progression of colorectal cancer is primarily mediated through metabolic reprogramming mechanisms.

Biotechnology2024arXiv
Periodicals

Dynamics of polymer ejection from capsid

Polymer ejection from a capsid through a nanoscale pore is an important biological process with relevance to modern biotechnology. Here, we study generic capsid ejection using Langevin dynamics. We show that even when the ejection takes place within the drift-dominated region there is a very high probability for the ejection process not to be completed. Introducing a small aligning force at the pore entrance enhances ejection dramatically. Such a pore asymmetry is a candidate for a mechanism by which a viral ejection is completed. By detailed high-resolution simulations we show that such capsid ejection is an out-of-equilibrium process that shares many common features with the much studied driven polymer translocation through a pore in a wall or a membrane. We find that the escape times scale with polymer length, $τ\sim N^α$. We show that for the pore without the asymmetry the previous predictions corroborated by Monte Carlo simulations do not hold. For the pore with the asymmetry the scaling exponent varies with the initial monomer density (monomers per capsid volume) $ρ$ inside the capsid. For very low densities $ρ\le 0.002$ the polymer is only weakly confined by the capsid, and we measure $α= 1.33$, which is close to $α= 1.4$ obtained for polymer translocation. At intermediate densities the scaling exponents $α= 1.25$ and $1.21$ for $ρ= 0.01$ and $0.02$, respectively. These scalings are in accord with a crude derivation for the lower limit $α= 1.2$. For the asymmetrical pore precise scaling breaks down, when the density exceeds the value for complete confinement by the capsid, $ρ\gtrapprox 0.25$. The high-resolution data show that the capsid ejection for both pores, analogously to polymer translocation, can be characterized as a multiplicative stochastic process that is dominated by small-scale transitions.

Biotechnology2013arXiv
Periodicals

MutaGAN: A Seq2seq GAN Framework to Predict Mutations of Evolving Protein Populations

The ability to predict the evolution of a pathogen would significantly improve the ability to control, prevent, and treat disease. Despite significant progress in other problem spaces, deep learning has yet to contribute to the issue of predicting mutations of evolving populations. To address this gap, we developed a novel machine learning framework using generative adversarial networks (GANs) with recurrent neural networks (RNNs) to accurately predict genetic mutations and evolution of future biological populations. Using a generalized time-reversible phylogenetic model of protein evolution with bootstrapped maximum likelihood tree estimation, we trained a sequence-to-sequence generator within an adversarial framework, named MutaGAN, to generate complete protein sequences augmented with possible mutations of future virus populations. Influenza virus sequences were identified as an ideal test case for this deep learning framework because it is a significant human pathogen with new strains emerging annually and global surveillance efforts have generated a large amount of publicly available data from the National Center for Biotechnology Information's (NCBI) Influenza Virus Resource (IVR). MutaGAN generated "child" sequences from a given "parent" protein sequence with a median Levenshtein distance of 2.00 amino acids. Additionally, the generator was able to augment the majority of parent proteins with at least one mutation identified within the global influenza virus population. These results demonstrate the power of the MutaGAN framework to aid in pathogen forecasting with implications for broad utility in evolutionary prediction for any protein population.

Biotechnology2020arXiv
Periodicals

A Conversation with Mike West

Mike West is currently the Arts & Sciences Distinguished Professor Emeritus of Statistics and Decision Sciences at Duke University. Mike's research in Bayesian analysis spans multiple interlinked areas: theory and methods of dynamic models in time series analysis, foundations of inference and decision analysis, multivariate and latent structure analysis, stochastic computation and optimisation, among others. Inter-disciplinary R&D has ranged across applications in commercial forecasting, dynamic networks, finance, econometrics, signal processing, climatology, systems biology, genomics and neuroscience, among other areas. Among Mike's currently active research areas are forecasting, causal prediction and decision analysis in business, economic policy and finance, as well as in personal decision making. Mike led the development of academic statistics at Duke University from 1990-2002, and has been broadly engaged in professional leadership elsewhere. He is past president of the International Society for Bayesian Analysis (ISBA), and has served in founding roles and as board member for several professional societies, national and international centres and institutes. Recipient of numerous awards, Mike has been active in research with various companies, banks, government agencies and academic centres, co-founder of a successful biotechnology company, and board member for several financial and IT companies. He has published 4 books, several edited volumes and over 200 papers. Mike has worked with many undergraduate and Master's research students, and as of 2025 has mentored around 65 primary PhD students and postdoctoral associates who moved to academic, industrial or governmental positions involving advanced statistical and data science research.

Biotechnology2025arXiv
Periodicals

The 2014 Magnetism Roadmap

Magnetism is a very fascinating and dynamic field. Especially in the last 30 years it has experienced many major advances in the full range from novel fundamental phenomena to new products. Applications such as hard disk drives and magnetic sensors are part of our daily life, and new applications, such as in non-volatile computer random access memory, are expected to surface shortly. Thus it is timely for describing the current status, and current and future challenges in the form of a Roadmap article. This 2014 Magnetism Roadmap provides a view on several selected, currently very active innovative developments. It consists of 12 sections, each written by an expert in the field and addressing a specific subject, with strong emphasize on future potential. This Roadmap cannot cover the entire field. We have selected several highly relevant areas without attempting to provide a full review - a future update will have room for more topics. The scope covers mostly nano-magnetic phenomena and applications, where surfaces and interfaces provide additional functionality. New developments in fundamental topics such as interacting nano-elements, novel magnon-based spintronics concepts, spin-orbit torques and spin-caloric phenomena are addressed. New materials, such as organic magnetic materials and permanent magnets are covered. New applications are presented such as nano-magnetic logic, non-local and domain-wall based devices, heat-assisted magnetic recording, magnetic random access memory, and applications in biotechnology. May the Roadmap serve as a guideline for future emerging research directions in modern magnetism.

Biotechnology2014arXiv
Periodicals

Matching Trace Element Distribution to Mineralogical Phases in Ancient Biotechnology-Derived Metallic Salts: a Multimodal Analysis

Conventional X-ray fluorescence (XRF) and X-ray diffraction (XRD) analysis applied to the investigation of ancient metal salts used as pigments and/or therapeutics provide bulk average compositions in two stand-alone data sets; however, major elements aside, these two sets cannot inform on the spatial distribution of one with respect to the other. To address this issue, we present here a multimodal approach incorporating spatially resolved XRF, XRD and nanoscale X-ray imaging applied to the analysis of archaeological and experimental samples of synthetic lead carbonate (PbCO3 - Greek psimythion); psimythion was used in antiquity as a cosmetic and/or a therapeutic for external applications. The experimental sample was produced according to a well-documented recipe dated to the 4th century BCE. In this paper we demonstrate that by using a multimodal approach we can confidently assign trace elements to individual crystalline or to infer the existence of non-crystalline phases. Although the assignment of an element to a phase (i.e. the location) is now possible, the origin underlying it (i.e. the mechanism) is not always clear. Trace elements do not 'control' chemical/mineralogical composition, but they can influence it. Our approach is particularly suited to following changes in the artefact's chemical/mineralogical profile, from its manufacture, to use and burial, to excavation and conservation.

Biotechnology2026arXiv
Periodicals

Interaction of Epithelial Cells with Surfaces and Surfaces Decorated by Molecules

A detailed understanding of the interface between living cells and substrate materials is of rising importance in many fields of medicine, biology and biotechnology. Cells at interfaces often form epithelia. The physical barrier that they form is one of their main functions. It is governed by the properties of the networks forming the cytoskeleton systems and by cell-to-cell contacts. Different substrates with varying surface properties modify the migration velocity of the cells. On the one hand one can change the materials composition. Organic and inorganic materials induce differing migration velocities in the same cell system. Within the same class of materials, a change of the surface stiffness or of the surface energy modifies the migration velocity, too. For our cell adhesion studies a variety of different, homogeneous substrates were used (polymers, bio-polymers, metals, oxides). In addition, an effective lithographic method, Polymer Blend Lithography (PBL), is reported, to produce patterned Self-Assembled Monolayers (SAM) on solid substrates featuring two or three different chemical functionalities. This we achieve without the use of conventional lithography like e-beam or UV lithography, only by using self-organization. These surfaces are decorated with a Teflon-like and with an amino-functionalized molecular layer. The resulting pattern is a copy of a previously created self-organized polymer pattern, featuring a scalable lateral domain size in the sub-micron range down below 100 nanometers. The resulting monolayer pattern features a high chemical and biofunctional contrast with feature sizes in the range of cell adhesion complexes like e.g. focal adhesion points.

Biotechnology2013arXiv
Periodicals

Temporal Dynamics of Microbial Communities in Anaerobic Digestion: Influence of Temperature and Feedstock Composition on Reactor Performance and Stability

Anaerobic digestion (AD) offers a sustainable biotechnology to recover resources from carbon-rich wastewater, such as food-processing wastewater. Despite crude wastewater characterisation, the impact of detailed chemical fingerprinting on AD remains underexplored. This study investigated the influence of fermentation-wastewater composition and operational parameters on AD over time to identify critical factors influencing reactor biodiversity and performance. Eighteen reactors were operated under various operational conditions using mycoprotein fermentation wastewater. Detailed chemical analysis fingerprinted the molecules in the fermentation wastewater throughout AD including sugars, sugar alcohols and volatile fatty acids (VFAs). Sequencing revealed distinct microbiome profiles linked to temperature and reactor configuration, with mesophilic conditions supporting a more diverse and densely connected microbiome. Significant elevations in Methanomassiliicoccus were correlated to high butyric acid concentrations and decreased biogas production, further elucidating the role of this newly discovered methanogen. Dissimilarity analysis demonstrated the importance of individual molecules on microbiome diversity, highlighting the need for detailed chemical fingerprinting in AD studies of microbial trends. Machine learning (ML) models predicting reactor performance achieved high accuracy based on operational parameters and microbial taxonomy. Operational parameters had the most substantial influence on chemical oxygen demand removal, whilst Oscillibacter and two Clostridium sp. were highlighted as key factors in biogas production. By integrating detailed chemical and biological fingerprinting with ML models this research presents a novel approach to advance our understanding of AD microbial ecology, offering insights for industrial applications of sustainable waste-to-energy systems.

Biotechnology2025arXiv
Periodicals

A Tutorial to Multirate Extended Kalman Filter Design for Monitoring of Agricultural Anaerobic Digestion Plants

In many applications of biotechnology, measurements are available at different sampling rates, e.g., due to online sensors and offline lab analysis. Offline measurements typically involve time delays that may be unknown a priori due to the underlying laboratory procedures. This multirate (MR) setting poses a challenge to Kalman filtering, where conventionally measurement data is assumed to be available on an equidistant time grid and without delays. This tutorial paper derives the MR version of an extended Kalman filter (EKF) based on sample state augmentation, and applies it to the anaerobic digestion (AD) process in a simulative agricultural setting. The performance of the MR-EKF is investigated for various scenarios including varying delay lengths, measurement noise levels, plant-model mismatch (PMM), and initial state error. Provided with an adequate tuning, the MR-EKF can reliably estimate the process state and, thus, appropriately fuse the delayed offline measurements and smooth the noisy online measurements. Because of the sample state augmentation approach, the delay length of offline measurements does not critically effect the performance of the state estimation, provided that observability is not lost during the delays. Poor state initialization and PMM affect convergence more than measurement noise levels. Furthermore, selecting an appropriate tuning was found to be critically important for successful application of the MR-EKF for which a systematic approach is presented. This tutorial provides implementation guidance for practitioners seeking to successfully apply state estimation for multirate systems. Thus, it contributes to the development of demand-driven operation of biogas plants, which may aid in stabilizing a renewable electricity grid.

Biotechnology2025arXiv
Periodicals

Systems of Global Governance in the Era of Human-Machine Convergence

Technology is increasingly shaping our social structures and is becoming a driving force in altering human biology. Besides, human activities already proved to have a significant impact on the Earth system which in turn generates complex feedback loops between social and ecological systems. Furthermore, since our species evolved relatively fast from small groups of hunter-gatherers to large and technology-intensive urban agglomerations, it is not a surprise that the major institutions of human society are no longer fit to cope with the present complexity. In this note we draw foundational parallelisms between neurophysiological systems and ICT-enabled social systems, discussing how frameworks rooted in biology and physics could provide heuristic value in the design of evolutionary systems relevant to politics and economics. In this regard we highlight how the governance of emerging technology (i.e. nanotechnology, biotechnology, information technology, and cognitive science), and the one of climate change both presently confront us with a number of connected challenges. In particular: historically high level of inequality; the co-existence of growing multipolar cultural systems in an unprecedentedly connected world; the unlikely reaching of the institutional agreements required to deviate abnormal trajectories of development. We argue that wise general solutions to such interrelated issues should embed the deep understanding of how to elicit mutual incentives in the socio-economic subsystems of Earth system in order to jointly concur to a global utility function (e.g. avoiding the reach of planetary boundaries and widespread social unrest). We leave some open questions on how techno-social systems can effectively learn and adapt with respect to our understanding of geopolitical complexity.

Biotechnology2018arXiv
Periodicals

Using association rule mining and ontologies to generate metadata recommendations from multiple biomedical databases

Metadata-the machine-readable descriptions of the data-are increasingly seen as crucial for describing the vast array of biomedical datasets that are currently being deposited in public repositories. While most public repositories have firm requirements that metadata must accompany submitted datasets, the quality of those metadata is generally very poor. A key problem is that the typical metadata acquisition process is onerous and time consuming, with little interactive guidance or assistance provided to users. Secondary problems include the lack of validation and sparse use of standardized terms or ontologies when authoring metadata. There is a pressing need for improvements to the metadata acquisition process that will help users to enter metadata quickly and accurately. In this paper we outline a recommendation system for metadata that aims to address this challenge. Our approach uses association rule mining to uncover hidden associations among metadata values and to represent them in the form of association rules. These rules are then used to present users with real-time recommendations when authoring metadata. The novelties of our method are that it is able to combine analyses of metadata from multiple repositories when generating recommendations and can enhance those recommendations by aligning them with ontology terms. We implemented our approach as a service integrated into the CEDAR Workbench metadata authoring platform, and evaluated it using metadata from two public biomedical repositories: US-based National Center for Biotechnology Information (NCBI) BioSample and European Bioinformatics Institute (EBI) BioSamples. The results show that our approach is able to use analyses of previous entered metadata coupled with ontology-based mappings to present users with accurate recommendations when authoring metadata.

Biotechnology2019arXiv
Periodicals

SA-GNAS: Seed Architecture Expansion for Efficient Large-scale Graph Neural Architecture Search

GNAS (Graph Neural Architecture Search) has demonstrated great effectiveness in automatically designing the optimal graph neural architectures for multiple downstream tasks, such as node classification and link prediction. However, most existing GNAS methods cannot efficiently handle large-scale graphs containing more than million-scale nodes and edges due to the expensive computational and memory overhead. To scale GNAS on large graphs while achieving better performance, we propose SA-GNAS, a novel framework based on seed architecture expansion for efficient large-scale GNAS. Similar to the cell expansion in biotechnology, we first construct a seed architecture and then expand the seed architecture iteratively. Specifically, we first propose a performance ranking consistency-based seed architecture selection method, which selects the architecture searched on the subgraph that best matches the original large-scale graph. Then, we propose an entropy minimization-based seed architecture expansion method to further improve the performance of the seed architecture. Extensive experimental results on five large-scale graphs demonstrate that the proposed SA-GNAS outperforms human-designed state-of-the-art GNN architectures and existing graph NAS methods. Moreover, SA-GNAS can significantly reduce the search time, showing better search efficiency. For the largest graph with billion edges, SA-GNAS can achieve 2.8 times speedup compared to the SOTA large-scale GNAS method GAUSS. Additionally, since SA-GNAS is inherently parallelized, the search efficiency can be further improved with more GPUs. SA-GNAS is available at https://github.com/PasaLab/SAGNAS.

Biotechnology2024arXiv
Periodicals

Trans-dimensional Bayesian model averaging for $^{13}$C-based metabolic flux analysis: Evidence-based flux inference under structural model uncertainty

Accurate quantification of intracellular metabolic fluxes is central to systems biology and biotechnology. Flux estimation relies on biochemical network models, with $^{13}$C metabolic flux analysis (MFA) being the state-of-the-art approach. However, isotope labeling data are often insufficient to uniquely support a single network formulation. In such cases, flux estimates become model-dependent, highlighting the need for methods that explicitly account for structural uncertainty. Bayesian model averaging (BMA) provides a principled framework for this purpose, but its application to $^{13}$C-MFA has so far been restricted to uncertainty in reaction bidirectionality within fixed network topologies. We introduce a scalable Bayesian inference framework for $^{13}$C-MFA, Bayesian model set averaging, that applies BMA to encompass uncertainty in reactions and pathways. Our approach combines reversible jump Markov chain Monte Carlo for trans-dimensional exploration of model spaces with diffusive nested sampling for robust estimation of model evidences, enabling averaging over large families of metabolic network models. Using illustrative and application-scale synthetic case studies, we demonstrate that the method yields robust flux estimates, reveals when multiple network configurations are statistically indistinguishable, and recovers data-supported model structures. Importantly, rather than committing to a single model, the framework manages structural uncertainty: under limited data, competing models are retained, whereas increasing data informativeness improved model and flux recovery. The approach scales to billions of model variants, providing a practical foundation for uncertainty- and misspecification-aware quantitative flux inference in $^{13}$C-MFA.

Biotechnology2026arXiv
Periodicals

Optimizing Agricultural Research: A RAG-Based Approach to Mycorrhizal Fungi Information

Retrieval-Augmented Generation (RAG) represents a transformative approach within natural language processing (NLP), combining neural information retrieval with generative language modeling to enhance both contextual accuracy and factual reliability of responses. Unlike conventional Large Language Models (LLMs), which are constrained by static training corpora, RAG-powered systems dynamically integrate domain-specific external knowledge sources, thereby overcoming temporal and disciplinary limitations. In this study, we present the design and evaluation of a RAG-enabled system tailored for Mycophyto, with a focus on advancing agricultural applications related to arbuscular mycorrhizal fungi (AMF). These fungi play a critical role in sustainable agriculture by enhancing nutrient acquisition, improving plant resilience under abiotic and biotic stresses, and contributing to soil health. Our system operationalizes a dual-layered strategy: (i) semantic retrieval and augmentation of domain-specific content from agronomy and biotechnology corpora using vector embeddings, and (ii) structured data extraction to capture predefined experimental metadata such as inoculation methods, spore densities, soil parameters, and yield outcomes. This hybrid approach ensures that generated responses are not only semantically aligned but also supported by structured experimental evidence. To support scalability, embeddings are stored in a high-performance vector database, allowing near real-time retrieval from an evolving literature base. Empirical evaluation demonstrates that the proposed pipeline retrieves and synthesizes highly relevant information regarding AMF interactions with crop systems, such as tomato (Solanum lycopersicum). The framework underscores the potential of AI-driven knowledge discovery to accelerate agroecological innovation and enhance decision-making in sustainable farming systems.

Biotechnology2025arXiv
Periodicals

Taec: a Manually annotated text dataset for trait and phenotype extraction and entity linking in wheat breeding literature

Wheat varieties show a large diversity of traits and phenotypes. Linking them to genetic variability is essential for shorter and more efficient wheat breeding programs. Newly desirable wheat variety traits include disease resistance to reduce pesticide use, adaptation to climate change, resistance to heat and drought stresses, or low gluten content of grains. Wheat breeding experiments are documented by a large body of scientific literature and observational data obtained in-field and under controlled conditions. The cross-referencing of complementary information from the literature and observational data is essential to the study of the genotype-phenotype relationship and to the improvement of wheat selection. The scientific literature on genetic marker-assisted selection describes much information about the genotype-phenotype relationship. However, the variety of expressions used to refer to traits and phenotype values in scientific articles is a hinder to finding information and cross-referencing it. When trained adequately by annotated examples, recent text mining methods perform highly in named entity recognition and linking in the scientific domain. While several corpora contain annotations of human and animal phenotypes, currently, no corpus is available for training and evaluating named entity recognition and entity-linking methods in plant phenotype literature. The Triticum aestivum trait Corpus is a new gold standard for traits and phenotypes of wheat. It consists of 540 PubMed references fully annotated for trait, phenotype, and species named entities using the Wheat Trait and Phenotype Ontology and the species taxonomy of the National Center for Biotechnology Information. A study of the performance of tools trained on the Triticum aestivum trait Corpus shows that the corpus is suitable for the training and evaluation of named entity recognition and linking.

Biotechnology2024arXiv
Periodicals

Toward Total Recall: Enhancing FAIRness through AI-Driven Metadata Standardization

Scientific metadata often suffer from incompleteness, inconsistency, and formatting errors, which hinder effective discovery and reuse of the associated datasets. We present a method that combines GPT-4 with structured metadata templates from the CEDAR knowledge base to automatically standardize metadata and to ensure compliance with established standards. A CEDAR template specifies the expected fields of a metadata submission and their permissible values. Our standardization process involves using CEDAR templates to guide GPT-4 in accurately correcting and refining metadata entries in bulk, resulting in significant improvements in metadata retrieval performance, especially in recall -- the proportion of relevant datasets retrieved from the total relevant datasets available. Using the BioSample and GEO repositories maintained by the National Center for Biotechnology Information (NCBI), we demonstrate that retrieval of datasets whose metadata are altered by GPT-4 when provided with CEDAR templates (GPT-4+CEDAR) is substantially better than retrieval of datasets whose metadata are in their original state and that of datasets whose metadata are altered using GPT-4 with only data-dictionary guidance (GPT-4+DD). The average recall increases dramatically, from 17.65\% with baseline raw metadata to 62.87\% with GPT-4+CEDAR. Furthermore, we evaluate the robustness of our approach by comparing GPT-4 against other large language models, including LLaMA-3 and MedLLaMA2, demonstrating consistent performance advantages for GPT-4+CEDAR. These results underscore the transformative potential of combining advanced language models with symbolic models of standardized metadata structures for more effective and reliable data retrieval, thus accelerating scientific discoveries and data-driven research.

Biotechnology2025arXiv
Periodicals

Single Cancer Cell Detection by Near Infrared Microspectroscopy, Infrared Chemical Imaging and Fluorescence Microspectroscopy

Novel techniques are currently being developed and established for the accurate chemical analysis and detection of single cancer cells, single embryos and single seeds by Fourier Transform Near Infrared (FT-NIR) Microspectroscopy, Fourier Transform Infrared (FT-IR), Fluorescence and High-Resolution NMR (HR-NMR). The first FT-NIR chemical images of biological systems approaching 1micron resolution are here reported. 400 and 500 MHz, H-1 NMR analyses were carried out that allowed the selection of mutagenized embryos. Detailed chemical analyses are being demonstrated to be also possible by FT-NIR Chemical Imaging/ Microspectroscopy of single cancer cells. FT-NIR Microspectroscopy and Chemical Imaging are also shown to be potentially important in Functional Genomics and Proteomics research through the rapid and accurate detection of high-content microarrays (HCMA). Multi-photon (MP), pulsed femtosecond laser NIR Fluorescence Excitation techniques were shown to be capable of Single Molecule Detection (SMD. Thus, MP NIR excitation for Fluorescence Correlation Spectroscopy (FCS) allowed not only single molecule detection, but also molecular dynamics observations and high resolution, submicron imaging of sub-femtoliter volumes inside living cells with 0.25 micron spatial resolution, in both normal and cancer cells, as well as neoplastic tissues. These novel, ultra-sensitive and rapid FT-NIR/FCS analyses have, therefore, substantial potential for numerous applications in important research areas, such as: medicine, medical/cancer research, pharmacology, agricultural biotechnology, food safety, as well as clinical diagnosis of viral diseases and cancers.

Biotechnology2004arXiv
Periodicals

Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation

Designing proteins de novo with tailored structural, physicochemical, and functional properties remains a grand challenge in biotechnology, medicine, and materials science, due to the vastness of sequence space and the complex coupling between sequence, structure, and function. Current state-of-the-art generative methods, such as protein language models (PLMs) and diffusion-based architectures, often require extensive fine-tuning, task-specific data, or model reconfiguration to support objective-directed design, thereby limiting their flexibility and scalability. To overcome these limitations, we present a decentralized, agent-based framework inspired by swarm intelligence for de novo protein design. In this approach, multiple large language model (LLM) agents operate in parallel, each assigned to a specific residue position. These agents iteratively propose context-aware mutations by integrating design objectives, local neighborhood interactions, and memory and feedback from previous iterations. This position-wise, decentralized coordination enables emergent design of diverse, well-defined sequences without reliance on motif scaffolds or multiple sequence alignments, validated with experiments on proteins with alpha helix and coil structures. Through analyses of residue conservation, structure-based metrics, and sequence convergence and embeddings, we demonstrate that the framework exhibits emergent behaviors and effective navigation of the protein fitness landscape. Our method achieves efficient, objective-directed designs within a few GPU-hours and operates entirely without fine-tuning or specialized training, offering a generalizable and adaptable solution for protein design. Beyond proteins, the approach lays the groundwork for collective LLM-driven design across biomolecular systems and other scientific discovery tasks.

Biotechnology2025arXiv
Periodicals

FoodMem: Near Real-time and Precise Food Video Segmentation

Food segmentation, including in videos, is vital for addressing real-world health, agriculture, and food biotechnology issues. Current limitations lead to inaccurate nutritional analysis, inefficient crop management, and suboptimal food processing, impacting food security and public health. Improving segmentation techniques can enhance dietary assessments, agricultural productivity, and the food production process. This study introduces the development of a robust framework for high-quality, near-real-time segmentation and tracking of food items in videos, using minimal hardware resources. We present FoodMem, a novel framework designed to segment food items from video sequences of 360-degree unbounded scenes. FoodMem can consistently generate masks of food portions in a video sequence, overcoming the limitations of existing semantic segmentation models, such as flickering and prohibitive inference speeds in video processing contexts. To address these issues, FoodMem leverages a two-phase solution: a transformer segmentation phase to create initial segmentation masks and a memory-based tracking phase to monitor food masks in complex scenes. Our framework outperforms current state-of-the-art food segmentation models, yielding superior performance across various conditions, such as camera angles, lighting, reflections, scene complexity, and food diversity. This results in reduced segmentation noise, elimination of artifacts, and completion of missing segments. Here, we also introduce a new annotated food dataset encompassing challenging scenarios absent in previous benchmarks. Extensive experiments conducted on MetaFood3D, Nutrition5k, and Vegetables & Fruits datasets demonstrate that FoodMem enhances the state-of-the-art by 2.5% mean average precision in food video segmentation and is 58 x faster on average.

Biotechnology2024arXiv
Periodicals

Prediction by Machine Learning Analysis of Genomic Data Phenotypic Frost Tolerance in Perccottus glenii

Analysis of the genome sequence of Perccottus glenii, the only fish known to possess freeze tolerance, holds significant importance for understanding how organisms adapt to extreme environments, Traditional biological analysis methods are time-consuming and have limited accuracy, To address these issues, we will employ machine learning techniques to analyze the gene sequences of Perccottus glenii, with Neodontobutis hainanens as a comparative group, Firstly, we have proposed five gene sequence vectorization methods and a method for handling ultra-long gene sequences, We conducted a comparative study on the three vectorization methods: ordinal encoding, One-Hot encoding, and K-mer encoding, to identify the optimal encoding method, Secondly, we constructed four classification models: Random Forest, LightGBM, XGBoost, and Decision Tree, The dataset used by these classification models was extracted from the National Center for Biotechnology Information database, and we vectorized the sequence matrices using the optimal encoding method, K-mer, The Random Forest model, which is the optimal model, achieved a classification accuracy of up to 99, 98 , Lastly, we utilized SHAP values to conduct an interpretable analysis of the optimal classification model, Through ten-fold cross-validation and the AUC metric, we identified the top 10 features that contribute the most to the model's classification accuracy, This demonstrates that machine learning methods can effectively replace traditional manual analysis in identifying genes associated with the freeze tolerance phenotype in Perccottus glenii.

Biotechnology2024arXiv
Periodicals

Metal-Enhanced Near-Infrared Fluorescence by Micropatterned Gold Nanocages

In metal-enhanced fluorescence (MEF), the localized surface plasmon resonances of metallic nanostructures amplify the absorption of excitation light and assist in radiating the consequent fluorescence of nearby molecules to the far-field. This effect is at the base of various technologies that have strong impact on fields such as optics, medical diagnostics and biotechnology. Among possible emission bands, those in the near-infrared (NIR) are particularly intriguing and widely used in proteomics and genomics due to its noninvasive character for biomolecules, living cells, and tissues, which greatly motivates the development of effective, and eventually multifunctional NIR-MEF platforms. Here we demonstrate NIR-MEF substrates based on Au nanocages micropatterned with a tight spatial control. The dependence of the fluorescence enhancement on the distance between the nanocage and the radiating dipoles is investigated experimentally and modeled by taking into account the local electric field enhancement and the modified radiation and absorption rates of the emitting molecules. At a distance around 80 nm, a maximum enhancement up to 2-7 times with respect to the emission from pristine dyes (in the region 660 nm-740 nm) is estimated for films and electrospun nanofibers. Due to their chemical stability, finely tunable plasmon resonances, and large light absorption cross sections, Au nanocages are ideal NIR-MEF agents. When these properties are integrated with the hollow interior and controllable surface porosity, it is feasible to develop a nanoscale system for targeted drug delivery with the diagnostic information encoded in the fluorophore.

Biotechnology2015arXiv