Academic Digital Library for institutions, students, and solo learners

Discovery

Content search and filters

This is the first search surface for the seven content types. Next we will connect full-text search and metadata-specific filters.

Results 1,892

Periodicals

Peptipedia: a comprehensive database for peptide research supported by Assembled predictive models and Data Mining approaches

Motivation: Peptides have attracted the attention in this century due to their remarkable therapeutic properties. Computational tools are being developed to take advantage of existing information, encapsulating knowledge and making it available in a simple way for general public use. However, these are property-specific redundant data systems, and usually do not display the data in a clear way. In some cases, information download is not even possible. This data needs to be available in a simple form for drug design and other biotechnological applications. Results: We developed Peptipedia, a user-friendly database and web application to search, characterise and analyse peptide sequences. Our tool integrates the information from thirty previously reported databases, making it the largest repository of peptides with recorded activities so far. Besides, we implemented a variety of services to increase our tool's usability. The significant differences of our tools with other existing alternatives becomes a substantial contribution to develop biotechnological and bioengineering applications for peptides. Availability: Peptipedia is available for non-commercial use as an open-access software, licensed under the GNU General Public License, version GPL 3.0. The web platform is publicly available at pesb2.cl/peptipedia. Both the source code and sample datasets are available in the GitHub repository https://github.com/CristoferQ/PeptideDatabase. Contact: david.medina@cebib.cl, ana.sanchez@ing.uchile.cl

Biotechnology2021arXiv
Periodicals

C3-Diff: Super-resolving Spatial Transcriptomics via Cross-modal Cross-content Contrastive Diffusion Modelling

The rapid advancement of spatial transcriptomics (ST), i.e., spatial gene expressions, has made it possible to measure gene expression within original tissue, enabling us to discover molecular mechanisms. However, current ST platforms frequently suffer from low resolution, limiting the in-depth understanding of spatial gene expression. Super-resolution approaches promise to enhance ST maps by integrating histology images with gene expressions of profiled tissue spots. However, it remains a challenge to model the interactions between histology images and gene expressions for effective ST enhancement. This study presents a cross-modal cross-content contrastive diffusion framework, called C3-Diff, for ST enhancement with histology images as guidance. In C3-Diff, we firstly analyze the deficiency of traditional contrastive learning paradigm, which is then refined to extract both modal-invariant and content-invariant features of ST maps and histology images. Further, to overcome the problem of low sequencing sensitivity in ST maps, we perform nosing-based information augmentation on the surface of feature unit hypersphere. Finally, we propose a dynamic cross-modal imputation-based training strategy to mitigate ST data scarcity. We tested C3-Diff by benchmarking its performance on four public datasets, where it achieves significant improvements over competing methods. Moreover, we evaluate C3-Diff on downstream tasks of cell type localization, gene expression correlation and single-cell-level gene expression prediction, promoting AI-enhanced biotechnology for biomedical research and clinical applications. Codes are available at https://github.com/XiaofeiWang2018/C3-Diff.

Biotechnology2025arXiv
Periodicals

Nanoscale Communication with Brownian Motion

In this paper, the problem of communicating using chemical messages propagating using Brownian motion, rather than electromagnetic messages propagating as waves in free space or along a wire, is considered. This problem is motivated by nanotechnological and biotechnological applications, where the energy cost of electromagnetic communication might be prohibitive. Models are given for communication using particles that propagate with Brownian motion, and achievable capacity results are given. Under conservative assumptions, it is shown that rates exceeding one bit per particle are achievable.

Biotechnology2007arXiv
Periodicals

Normalized topological indices discriminate between architectures of branched macromolecules

Branching architecture characterizes numerous systems, ranging from synthetic (hyper)branched polymers and biomolecules such as lignin, amylopectin, and nucleic acids to tracheal and neuronal networks. Its ubiquity reflects the many favourable properties that arise because of it. For instance, branched macromolecules are spatially compact and have a high surface functionality, which impacts their phase characteristics and self-assembly behaviour, among others. The relationship between branching and physical properties has been studied by mapping macromolecules to mathematical trees whose architecture can be characterized using topological indices. These indices, however, do not allow for a comparison of macromolecules that map to trees of different size, be it due to different mapping procedures or differences in their molecular weight. To alleviate this, we introduce a novel normalization of topological indices using estimates of their probability density functions. We determine two optimal normalized topological indices and construct a phase space that enables a robust discrimination between different architectures of branched macromolecules. We demonstrate the necessity of such a phase space on two practical applications, one being ribonucleic acid (RNA) molecules with various branching topologies and the other different methods of coarse-graining branched macromolecules. Our approach can be applied to any type of branched molecules and extended as needed to other topological indices, making it useful across a wide range of fields where branched molecules play an important role, including polymer physics, green chemistry, bioengineering, biotechnology, and medicine.

Biotechnology2024arXiv
Periodicals

Switching-time bioprocess control with pulse-width-modulated optogenetics

Biotechnology can benefit from dynamic control to improve production efficiency. In this context, optogenetics enables modulation of gene expression using light as an external input, allowing fine-tuning of protein levels to unlock dynamic metabolic control and regulation of cell growth. Optogenetic systems can be actuated by light intensity. However, relying solely on intensity-driven control (i.e., signal amplitude) may fail to properly tune optogenetic bioprocesses when the dose-response relationship (i.e., light intensity versus gene-expression strength) is steep. In these cases, tunability is effectively constrained to either fully active or fully repressed gene expression, with little intermediate regulation. Pulse-width modulation can alleviate this issue by alternating between fully ON and OFF light intensity within forcing periods, thereby smoothing the average response and enhancing process controllability. Optimizing pulse-width-modulated optogenetics entails a switching-time optimal control problem with a binary input over multiple forcing periods. While this can be formulated as a mixed-integer optimization problem on a refined control grid with monotonic input constraints, the number of decision variables can grow rapidly with increasing control-grid resolution within forcing periods and with the total number of forcing periods, complicating the task. Here, we propose an alternative solution based on reinforcement learning. We parametrize control actions via the duty cycle, a continuous proxy variable that encodes the ON-to-OFF switching time within each forcing period, thereby respecting the intrinsic binary nature of the light intensity while avoiding fine-grid binary decision variables.

Biotechnology2025arXiv
Periodicals

Machine Learning Modeling Of SiRNA Structure-Potency Relationship With Applications Against Sars-Cov-2 Spike Gene

The pharmaceutical Research and development (R&D) process is lengthy and costly, taking nearly a decade to bring a new drug to the market. However, advancements in biotechnology, computational methods, and machine learning algorithms have the potential to revolutionize drug discovery, speeding up the process and improving patient outcomes. The COVID-19 pandemic has further accelerated and deepened the recognition of the potential of these techniques, especially in the areas of drug repurposing and efficacy predictions. Meanwhile, non-small molecule therapeutic modalities such as cell therapies, monoclonal antibodies, and RNA interference (RNAi) technology have gained importance due to their ability to target specific disease pathways and/or patient populations. In the field of RNAi, many experiments have been carried out to design and select highly efficient siRNAs. However, the established patterns for efficient siRNAs are sometimes contradictory and unable to consistently determine the most potent siRNA molecules against a target mRNA. Thus, this paper focuses on developing machine learning models based on the cheminformatics representation of the nucleotide composition (i.e. AUTGC) of siRNA to predict their potency and aid the selection of the most efficient siRNAs for further development. The PLS (Partial Least Square) and SVR (Support Vector Regression) machine learning models built in this work outperformed previously published models. These models can help in predicting siRNA potency and aid in selecting the best siRNA molecules for experimental validation and further clinical development. The study has demonstrated the potential of AI/machine learning models to help expedite siRNA-based drug discovery including the discovery of potent siRNAs against SARS-CoV-2.

Biotechnology2024arXiv
Periodicals

Democratising Artificial Intelligence for Pandemic Preparedness and Global Governance in Latin American and Caribbean Countries

Infectious diseases, transmitted directly or indirectly, are among the leading causes of epidemics and pandemics. Consequently, several open challenges exist in predicting epidemic outbreaks, detecting variants, tracing contacts, discovering new drugs, and fighting misinformation. Artificial Intelligence (AI) can provide tools to deal with these scenarios, demonstrating promising results in the fight against the COVID-19 pandemic. AI is becoming increasingly integrated into various aspects of society. However, ensuring that AI benefits are distributed equitably and that they are used responsibly is crucial. Multiple countries are creating regulations to address these concerns, but the borderless nature of AI requires global cooperation to define regulatory and guideline consensus. Considering this, The Global South AI for Pandemic & Epidemic Preparedness & Response Network (AI4PEP) has developed an initiative comprising 16 projects across 16 countries in the Global South, seeking to strengthen equitable and responsive public health systems that leverage Southern-led responsible AI solutions to improve prevention, preparedness, and response to emerging and re-emerging infectious disease outbreaks. This opinion introduces our branches in Latin American and Caribbean (LAC) countries and discusses AI governance in LAC in the light of biotechnology. Our network in LAC has high potential to help fight infectious diseases, particularly in low- and middle-income countries, generating opportunities for the widespread use of AI techniques to improve the health and well-being of their communities.

Biotechnology2024arXiv
Periodicals

Comprehensive review of models and methods for inferences in bio-chemical reaction networks

Key processes in biological and chemical systems are described by networks of chemical reactions. From molecular biology to biotechnology applications, computational models of reaction networks are used extensively to elucidate their non-linear dynamics. Model dynamics are crucially dependent on parameter values which are often estimated from observations. Over past decade, the interest in parameter and state estimation in models of (bio-)chemical reaction networks (BRNs) grew considerably. Statistical inference problems are also encountered in many other tasks including model calibration, discrimination, identifiability and checking as well as optimum experiment design, sensitivity analysis, bifurcation analysis and other. The aim of this review paper is to explore developments of past decade to understand what BRN models are commonly used in literature, and for what inference tasks and inference methods. Initial collection of about 700 publications excluding books in computational biology and chemistry were screened to select over 260 research papers and 20 graduate theses concerning estimation problems in BRNs. The paper selection was performed as text mining using scripts to automate search for relevant keywords and terms. The outcome are tables revealing the level of interest in different inference tasks and methods for given models in literature as well as recent trends. In addition, a brief survey of general estimation strategies is provided to facilitate understanding of estimation methods which are used for BRNs. Our findings indicate that many combinations of models, tasks and methods are still relatively sparse representing new research opportunities to explore those that have not been considered - perhaps for a good reason. The paper concludes by discussing future research directions including research problems which cannot be directly deduced from presented tables.

Biotechnology2019arXiv
Periodicals

Facile preparation of agarose-chitosan hybrid materials and nanocomposite ionogels using an ionic liquid via dissolution, regeneration and sol-gel transition

We report simultaneous dissolution of agarose (AG) and chitosan (CH) in varying proportions in an ionic liquid (IL), 1-butyl-3-methylimidazolium chloride [C4mim][Cl]. Composite materials were constructed from AG-CH-IL solutions using the antisolvent methanol, and IL was recovered from the solutions. Composite materials could be uniformly decorated with silver oxide (Ag2O) nanoparticles (Ag NPs) to form nanocomposites in a single step by in situ synthesis of Ag NPs in AG-CH-IL sols, wherein the biopolymer moiety acted as both reducing and stabilizing agent. Cooling of Ag NPs-AG-CH-IL sols to room temperature resulted in high conductivity and high mechanical strength nanocomposite ionogels. The structure, stability and physiochemical properties of composite materials and nanocomposites were characterized by several analytical techniques, such as Fourier transform infrared (FTIR), CD spectroscopy, differential scanning colorimetric (DSC), thermogravimetric analysis (TGA), gel permeation chromatography (GPC), and scanning electron micrography (SEM). The result shows that composite materials have good thermal and conformational stability, compatibility and strong hydrogen bonding interactions between AG-CH complexes. Decoration of Ag NPs in composites and ionogels was confirmed by UV-Vis spectroscopy, SEM, TEM, EDAX and XRD. The mechanical and conducting properties of composite ionogels have been characterized by rheology and current-voltage measurements. Since Ag NPs show good antimicrobial activity, Ag NPs -AG-CH composite materials have the potential to be used in biotechnology and biomedical applications whereas nanocomposite ionogels will be suitable as precursors for applications such as quasi-solid dye sensitized solar cells, actuators, sensors or electrochromic displays.

Biotechnology2014arXiv
Periodicals

Accelerating Returns and the Qualitative Engine for Science

Ray Kurzweil described a thesis of accelerating returns, which is the most influential narratives in discussions of technological progress. Its central claim is that advances in multiple technological fields, especially compute, artificial intelligence, brain science, and biotechnology, interact in such a way that progress becomes self-amplifying and approximately exponential. This paper gives a simple mathematical interpretation of that claim and then argues that, even if such acceleration is real, it does not by itself resolve the central problem of scientific discovery. The reason is that accelerating returns apply most naturally to executional and infrastructural capability, whereas genuine discovery often depends on a different capacity: qualitative reasoning about when a current framework is structurally inadequate and what conceptual move is needed next. Recent ARC-AGI-3 results sharpen this distinction: humans solve the benchmark at ceiling, whereas frontier AI systems remain below 1%, indicating that the gap between current AI and human flexible reasoning is still very large. At the same time, Demis Hassabis has emphasized that humans must retain their sense of meaning and what they choose to focus their lives on, a reminder that the future of AI is not only a technical forecast but also a question of what forms of human understanding are worth preserving and transmitting. This paper positions the Qualitative Engine for Science (QES) [3] as a response to that missing capacity. In this view, the Kurzweil theory helps explain why quantitative capability may accelerate, while QES addresses the central problem in scientific discovery that acceleration alone does not solve. Its value does not depend on when AGI arrives, but on the fact that the processes of scientific discovery themselves constitute a form of human wisdom worth preserving, organizing, and making accessible.

Biotechnology2026arXiv
Periodicals

Perspectives for self-driving labs in synthetic biology

Self-driving labs (SDLs) combine fully automated experiments with artificial intelligence (AI) that decides the next set of experiments. Taken to their ultimate expression, SDLs could usher a new paradigm of scientific research, where the world is probed, interpreted, and explained by machines for human benefit. While there are functioning SDLs in the fields of chemistry and materials science, we contend that synthetic biology provides a unique opportunity since the genome provides a single target for affecting the incredibly wide repertoire of biological cell behavior. However, the level of investment required for the creation of biological SDLs is only warranted if directed towards solving difficult and enabling biological questions. Here, we discuss challenges and opportunities in creating SDLs for synthetic biology.

Biotechnology2022arXiv
Periodicals

Biomaterials: A trendy source to engineer functional entities -- An overview

The biomaterials exploitation in a sophisticated manner can provide extensive opportunities for experimentation in the field of interdisciplinary and multidisciplinary scientific research. Owing to the unique features of this trendy area, research scientists have been directed/redirected their interests in bio-based biomaterials for targeted applications in different sectors of the modern world. The present manuscript highlights the novel perspectives of biomaterials as a trendy source to engineer functional entities in numerous geometries for pharmaceuticals, cosmeceuticals, nutraceuticals, and other biotechnological or biomedical applications.

Biotechnology2018arXiv
Periodicals

Fourier Representations for Black-Box Optimization over Categorical Variables

Optimization of real-world black-box functions defined over purely categorical variables is an active area of research. In particular, optimization and design of biological sequences with specific functional or structural properties have a profound impact in medicine, materials science, and biotechnology. Standalone search algorithms, such as simulated annealing (SA) and Monte Carlo tree search (MCTS), are typically used for such optimization problems. In order to improve the performance and sample efficiency of such algorithms, we propose to use existing methods in conjunction with a surrogate model for the black-box evaluations over purely categorical variables. To this end, we present two different representations, a group-theoretic Fourier expansion and an abridged one-hot encoded Boolean Fourier expansion. To learn such representations, we consider two different settings to update our surrogate model. First, we utilize an adversarial online regression setting where Fourier characters of each representation are considered as experts and their respective coefficients are updated via an exponential weight update rule each time the black box is evaluated. Second, we consider a Bayesian setting where queries are selected via Thompson sampling and the posterior is updated via a sparse Bayesian regression model (over our proposed representation) with a regularized horseshoe prior. Numerical experiments over synthetic benchmarks as well as real-world RNA sequence optimization and design problems demonstrate the representational power of the proposed methods, which achieve competitive or superior performance compared to state-of-the-art counterparts, while improving the computation cost and/or sample efficiency, substantially.

Biotechnology2022arXiv
Periodicals

Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity

The rapid adoption of generative artificial intelligence (GenAI) in the biosciences is transforming biotechnology, medicine, and synthetic biology. Yet this advancement is intrinsically linked to new vulnerabilities, as GenAI lowers the barrier to misuse and introduces novel biosecurity threats, such as generating synthetic viral proteins or toxins. These dual-use risks are often overlooked, as existing safety guardrails remain fragile and can be circumvented through deceptive prompts or jailbreak techniques. In this Perspective, we first outline the current state of GenAI in the biosciences and emerging threat vectors ranging from jailbreak attacks and privacy risks to the dual-use challenges posed by autonomous AI agents. We then examine urgent gaps in regulation and oversight, drawing on insights from 130 expert interviews across academia, government, industry, and policy. A large majority ($\approx 76$\%) expressed concern over AI misuse in biology, and 74\% called for the development of new governance frameworks. Finally, we explore technical pathways to mitigation, advocating a multi-layered approach to GenAI safety. These defenses include rigorous data filtering, alignment with ethical principles during development, and real-time monitoring to block harmful requests. Together, these strategies provide a blueprint for embedding security throughout the GenAI lifecycle. As GenAI becomes integrated into the biosciences, safeguarding this frontier requires an immediate commitment to both adaptive governance and secure-by-design technologies.

Biotechnology2025arXiv
Periodicals

Data Augmentation Scheme for Raman Spectra with Highly Correlated Annotations

In biotechnology Raman Spectroscopy is rapidly gaining popularity as a process analytical technology (PAT) that measures cell densities, substrate- and product concentrations. As it records vibrational modes of molecules it provides that information non-invasively in a single spectrum. Typically, partial least squares (PLS) is the model of choice to infer information about variables of interest from the spectra. However, biological processes are known for their complexity where convolutional neural networks (CNN) present a powerful alternative. They can handle non-Gaussian noise and account for beam misalignment, pixel malfunctions or the presence of additional substances. However, they require a lot of data during model training, and they pick up non-linear dependencies in the process variables. In this work, we exploit the additive nature of spectra in order to generate additional data points from a given dataset that have statistically independent labels so that a network trained on such data exhibits low correlations between the model predictions. We show that training a CNN on these generated data points improves the performance on datasets where the annotations do not bear the same correlation as the dataset that was used for model training. This data augmentation technique enables us to reuse spectra as training data for new contexts that exhibit different correlations. The additional data allows for building a better and more robust model. This is of interest in scenarios where large amounts of historical data are available but are currently not used for model training. We demonstrate the capabilities of the proposed method using synthetic spectra of Ralstonia eutropha batch cultivations to monitor substrate, biomass and polyhydroxyalkanoate (PHA) biopolymer concentrations during of the experiments.

Biotechnology2024arXiv
Periodicals

Towards a modeling, optimization and predictive control framework for fed-batch metabolic cybergenetics

Biotechnology offers many opportunities for the sustainable manufacturing of valuable products. The toolbox to optimize bioprocesses includes \textit{extracellular} process elements such as the bioreactor design and mode of operation, medium formulation, culture conditions, feeding rates, etc. However, these elements are frequently insufficient for achieving optimal process performance or precise product composition. One can use metabolic and genetic engineering methods for optimization at the intracellular level. Nevertheless, those are often of static nature, failing when applied to dynamic processes or if disturbances occur. Furthermore, many bioprocesses are optimized empirically and implemented with little-to-no feedback control to counteract disturbances. The concept of cybergenetics has opened new possibilities to optimize bioprocesses by enabling online modulation of the gene expression of metabolism-relevant proteins via external inputs (e.g., light intensity in optogenetics). Here, we fuse cybergenetics with model-based optimization and predictive control for optimizing dynamic bioprocesses. To do so, we propose to use dynamic constraint-based models that integrate the dynamics of metabolic reactions, resource allocation, and inducible gene expression. We formulate a model-based optimal control problem to find the optimal process inputs. Furthermore, we propose using model predictive control to address uncertainties via online feedback. We focus on fed-batch processes, where the substrate feeding rate is an additional optimization variable. As a simulation example, we show the optogenetic control of the ATPase enzyme complex for dynamic modulation of enforced ATP wasting to adjust product yield and productivity.

Biotechnology2023arXiv
Periodicals

Reference environments: A universal tool for reproducibility in computational biology

The drive for reproducibility in the computational sciences has provoked discussion and effort across a broad range of perspectives: technological, legislative/policy, education, and publishing. Discussion on these topics is not new, but the need to adopt standards for reproducibility of claims made based on computational results is now clear to researchers, publishers and policymakers alike. Many technologies exist to support and promote reproduction of computational results: containerisation tools like Docker, literate programming approaches such as Sweave, knitr, iPython or cloud environments like Amazon Web Services. But these technologies are tied to specific programming languages (e.g. Sweave/knitr to R; iPython to Python) or to platforms (e.g. Docker for 64-bit Linux environments only). To date, no single approach is able to span the broad range of technologies and platforms represented in computational biology and biotechnology. To enable reproducibility across computational biology, we demonstrate an approach and provide a set of tools that is suitable for all computational work and is not tied to a particular programming language or platform. We present published examples from a series of papers in different areas of computational biology, spanning the major languages and technologies in the field (Python/R/MATLAB/Fortran/C/Java). Our approach produces a transparent and flexible process for replication and recomputation of results. Ultimately, its most valuable aspect is the decoupling of methods in computational biology from their implementation. Separating the 'how' (method) of a publication from the 'where' (implementation) promotes genuinely open science and benefits the scientific community as a whole.

Biotechnology2018arXiv
Periodicals

IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEads

The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2-3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silicomethodologies need to be improved to better select lead compounds that can proceed to later stages of the drug discovery protocol accelerating the entire process. No single methodological approach can achieve the necessary accuracy with required efficiency. Here we describe multiple algorithmic innovations to overcome this fundamental limitation, development and deployment of computational infrastructure at scale integrates multiple artificial intelligence and simulation-based approaches. Three measures of performance are:(i) throughput, the number of ligands per unit time; (ii) scientific performance, the number of effective ligands sampled per unit time and (iii) peak performance, in flop/s. The capabilities outlined here have been used in production for several months as the workhorse of the computational infrastructure to support the capabilities of the US-DOE National Virtual Biotechnology Laboratory in combination with resources from the EU Centre of Excellence in Computational Biomedicine.

Biotechnology2020arXiv
Periodicals

Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study

Large language models (LLMs) produce context inconsistency hallucinations, which are LLM generated outputs that are misaligned with the user prompt. This research project investigates whether prompt engineering (PE) methods can mitigate context inconsistency hallucinations in zero-shot LLM summarisation of scientific texts, where zero-shot indicates that the LLM relies purely on its pre-training data. Across eight yeast biotechnology research paper abstracts, six instruction-tuned LLMs were prompted with seven methods: a baseline prompt, two levels of increasing instruction complexity (PE-1 and PE-2), two levels of context repetition (CR-K1 and CR-K2), and two levels of random addition (RA-K1 and RA-K2). Context repetition involved the identification and repetition of K key sentences from the abstract, whereas random addition involved the repetition of K randomly selected sentences from the abstract, where K is 1 or 2. A total of 336 LLM-generated summaries were evaluated using six metrics: ROUGE-1, ROUGE-2, ROUGE-L, BERTScore, METEOR, and cosine similarity, which were used to compute the lexical and semantic alignment between the summaries and the abstracts. Four hypotheses on the effects of prompt methods on summary alignment with the reference text were tested. Statistical analysis on 3744 collected datapoints was performed using bias-corrected and accelerated (BCa) bootstrap confidence intervals and Wilcoxon signed-rank tests with Bonferroni-Holm correction. The results demonstrated that CR and RA significantly improve the lexical alignment of LLM-generated summaries with the abstracts. These findings indicate that prompt engineering has the potential to impact hallucinations in zero-shot scientific summarisation tasks.

Biotechnology2025arXiv
Periodicals

Multidisciplinary Cognitive Content of Nanoscience and Nanotechnology

This article examines the cognitive evolution and disciplinary diversity of nanotechnology as expressed through the terminology used in titles of nano journal articles. The analysis is based on the NanoBank bibliographic database of 287,106 nano articles published between 1981 and 2004. We perform multifaceted analyses of title words, focusing on 100 most frequent terms. Hierarchical clustering of title terms reveals three distinct time periods of cognitive development of nano research: formative (1981-1990), early (1991-1998), and current (after 1998). Early period is characterized by the introduction of thin film deposition techniques, while the current period is characterized by the increased focus on carbon nanotube and nanoparticle research. We introduce a method to identify disciplinary components of nanotechnology. It shows that the nano research is being carried out in a number of diverse parent disciplines. Currently only 5% of articles are published in dedicated nano-only journals. We find that some 85% of nano research today is multidisciplinary. Hierarchical clustering of disciplinary components reveals that the cognitive content of current nanoscience can be divided into nine clusters. Some clusters account for a large fraction of nano research and are identified with such parent disciplines as the condensed matter and applied physics, materials science, and analytical chemistry. Other clusters represent much smaller parts of nano research, but are as cognitively distinct. In the decreasing order of size, these fields are: polymer science, biotechnology, general chemistry, surface science, and pharmacology. Cognitive content of research published in nano-only journals is closest to nano research published in condensed matter and applied physics journals.

Biotechnology2012arXiv
Periodicals

SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions

Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of SciSchema.org, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology & Biotechnology, Materials & Chemistry, Imaging & Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.

Biotechnology2026arXiv
Periodicals

Leveraging partial coherence in interferometric microscopy to enhance nanoparticle detection sensitivity and throughput

Interferometric-based microscopies stand as powerful label-free approaches for monitoring and characterising chemical reactions and heterogeneous nanoparticle systems in real time with single particle sensitivity. Nevertheless, coherent artifacts, such as speckle and parasitic interferences, together with limited photon fluxes from spatially incoherent sources, pose an ongoing challenge in achieving both high sensitivity and throughput. In this study, we systematically characterise how partial coherence affects both the signal contrast and the background noise level; thus, it offers a route to improve the signal-to-noise ratio from single nanoparticles (NPs), irrespective of their size and composition; or the light source used. We first validate that lasers can be modified into partially coherent sources with performance matching that of spatially incoherent ones; while providing higher photon fluxes. Secondly, we demonstrate that tuning the degree of partial coherence not only enhances the detection sensitivity of both synthetic and biological NPs, but also affects how signal contrasts vary as a function of the focus position. Finally, we apply our findings to single-protein detection, confirming that these principles extend to differential imaging modalities, which deliver the highest sensitivity. Our results address a critical milestone in the detection of weakly scattering NPs in complex matrices, with wide-ranging applications in biotechnology, nanotechnology, chemical synthesis, and biosensing; ushering a new generation of microscopes that push both the sensitivity and throughput boundaries without requiring beam scanning.

Biotechnology2025arXiv
Periodicals

BioInfoBase : A Bioinformatics Resourceome

Over the past decade there has been a significant growth in bioinformatics databases, tools and resources. Although, bioinformatics is becoming more specific, increasing the number of bioinformatics-wares has made it difficult for researchers to find the most appropriate databases, tools or methods which match their needs. Our coordinated effort has been planned to establish a reference website in Bioinformatics as a public repository of tools, databases, directories and resources annotated with contextual information and organized by functional relevance. Within the first phase of BioInfoBase development, 22 experts in different fields of molecular biology contributed and more than 2500 records were registered, which are increasing daily. For each record submitted to the database of website almost all related data (40 features) has been extracted. These include information from the biological category and subcategory to the scientific article and developer information. Searching the query keyword(s) returns links containing the entered keyword(s) found within the different features of the records with more weights on the title, abstract and application fields. The search results simply provide the users with the most informative features of the records to select the most suitable ones. The usefulness of the returned results is ranked according to the matching score based on the Term Frequency-Inverse Document Frequency (TF-IDF) methods. Therefore, this search engine will screen a comprehensive index of bioinformatics tools, databases and resources and provide the best suited records (links) to the researchers need. The BioInfoBase resource is available at www.bioinfobase.info.

Biotechnology2016arXiv
Periodicals

Economic returns of research: the Pareto law and its implications

At what level should government or companies support research? This complex multi-faceted question encompasses such qualitative bonus as satisfying natural human curiosity, the quest for knowledge and the impact on education and culture, but one of its most scrutinized component reduces to the assessment of economic performance and wealth creation derived from research. In certain areas such as biotechnology, semi-conductor physics, optical communications, the impact of basic research is direct while, in other disciplines, the path from discovery to applications is full of surprises. As a consequence, there are persistent uncertainties in the quantification of the exact economic returns of public expenditure on basic research. Here, we suggest that these uncertainties have a fundamental origin to be found in the interplay between the intrinsic ``fat tail'' power law nature of the distribution of economic returns, characterized by a mathematically diverging variance, and the stochastic character of discovery rates. In the regime where the cumulative economic wealth derived from research is expected to exhibit a long-term positive trend, we show that strong fluctuations blur out significantly the short-time scales: a few major unpredictable innovations may provide a finite fraction of the total creation of wealth. In such a scenario, any attempt to assess the economic impact of research over a finite time horizon encompassing only a small number of major discoveries is bound to be highly unreliable. New tools, developed in the theory of self-similar and complex systems to tackle similar extreme fluctuations in Nature can be adapted to measure the economic benefits of research, which is intimately associated to this large variability.

Biotechnology1998arXiv