Academic Digital Library for institutions, students, and solo learners

Discovery

Content search and filters

This is the first search surface for the seven content types. Next we will connect full-text search and metadata-specific filters.

Results 1,892

Periodicals

From Explainable to Interpretable Deep Learning for Natural Language Processing in Healthcare: How Far from Reality?

Deep learning (DL) has substantially enhanced natural language processing (NLP) in healthcare research. However, the increasing complexity of DL-based NLP necessitates transparent model interpretability, or at least explainability, for reliable decision-making. This work presents a thorough scoping review of explainable and interpretable DL in healthcare NLP. The term "eXplainable and Interpretable Artificial Intelligence" (XIAI) is introduced to distinguish XAI from IAI. Different models are further categorized based on their functionality (model-, input-, output-based) and scope (local, global). Our analysis shows that attention mechanisms are the most prevalent emerging IAI technique. The use of IAI is growing, distinguishing it from XAI. The major challenges identified are that most XIAI does not explore "global" modelling processes, the lack of best practices, and the lack of systematic evaluation and benchmarks. One important opportunity is to use attention mechanisms to enhance multi-modal XIAI for personalized medicine. Additionally, combining DL with causal logic holds promise. Our discussion encourages the integration of XIAI in Large Language Models (LLMs) and domain-specific smaller models. In conclusion, XIAI adoption in healthcare requires dedicated in-house expertise. Collaboration with domain experts, end-users, and policymakers can lead to ready-to-use XIAI methods across NLP and medical tasks. While challenges exist, XIAI techniques offer a valuable foundation for interpretable NLP algorithms in healthcare.

Biotechnology2024arXiv
Periodicals

Enumeration of saturated and unsaturated substituted N-heterocycles

Mathematical and computational approaches in chemistry and biochemistry fill a gap in respect to the analysis of the physicochemical features of compounds and their functionality and provide an overview of known as well as yet unknown, but hypothetically possible structures. Nitrogen-containing heterocycles such as aziridine, azetidine and pyrrolidine bear a high potential in pharmacology, biotechnology and synthetic biology. Here, we present a mathematical enumeration procedure for all possible azaheterocycles with at least one substituent depending on the number of atoms in the ring, in the sense of saturated and unsaturated congeners. One subgroup belonging to that substance class is constituted by ring-shaped amino acids with a secondary amino group, such as proline. A recursion formula is derived, which results in a modified Lucas number series. Moreover, an explicit formula for determining the number of such substances based on the Golden Ratio is given and a second one, based on binomial coefficients, is newly derived. This enumeration is a helpful tool for construction or complementation of virtual compound databases and for computer-assisted chemical synthesis route planning.

Biotechnology2023arXiv
Periodicals

Isotopic Resonance Hypothesis: Experimental Verification by Escherichia coli Growth Measurements

Isotopic composition of reactants affects the rates of chemical and biochemical reactions. As a rule, enrichment of heavy stable isotopes leads to slower reactions. But the recent isotopic resonance hypothesis suggests that the dependence of the reaction rate upon the enrichment degree is not monotonous; instead, at some resonance isotopic compositions, the kinetics increases, while at off resonance compositions the same reactions progress slower. To test the predictions of this hypothesis for the elements C, H, N and O, we designed a precise (standard error plus or minus 0.05%) experiment to measure the bacterial growth parameters in minimal media with varying isotopic compositions. A number of predicted resonance conditions were tested, which kinetic enhancements as strong as plus 3% discovered at these conditions. The combined evidence extremely strongly supports the existence of isotopic resonances. This phenomenon has numerous implications for the origin of life and astrobiology, and possible applications in agriculture, biotechnology, medicine and other areas.

Biotechnology2014arXiv
Periodicals

Snapshot hyperspectral imaging with quantum correlated photons

Hyperspectral imaging (HSI) has a wide range of applications from environmental monitoring to biotechnology. Current snapshot HSI techniques all require a trade-off between spatial and spectral resolution and are thus unable to achieve high resolutions in both simultaneously. Additionally, the techniques are resource inefficient with most of the photons lost through spectral filtering. Here, we demonstrate a snapshot HSI technique utilizing the strong spectro-temporal correlations inherent in entangled photons using a modified quantum ghost spectroscopy system, where the target is directly imaged with one photon and the spectral information gained through ghost spectroscopy from the partner photon. As only a few rows of pixels near the edge of the camera are used for the spectrometer, almost no spatial resolution is sacrificed for spectral. Also since no spectral filtering is required, all photons contribute to the HSI process making the technique much more resource efficient.

Biotechnology2022arXiv
Periodicals

A note on the minimax solution for the two-stage group testing problem

Group testing is an active area of current research and has important applications in medicine, biotechnology, genetics, and product testing. There have been recent advances in design and estimation, but the simple Dorfman procedure introduced by R. Dorfman in 1943 is widely used in practice. In many practical situations the exact value of the probability p of being affected is unknown. We present both minimax and Bayesian solutions for the group size problem when p is unknown. For unbounded p we show that the minimax solution for group size is 8, while using a Bayesian strategy with Jeffreys prior results in a group size of 13. We also present solutions when p is bounded from above. For the practitioner we propose strong justification for using a group size of between eight to thirteen when a constraint on p is not incorporated and provide useable code for computing the minimax group size under a constrained p.

Biotechnology2014arXiv
Periodicals

RAPTOR: Ravenous Throughput Computing

We describe the design, implementation and performance of the RADICAL-Pilot task overlay (RAPTOR). RAPTOR enables the execution of heterogeneous tasks -- i.e., functions and executables with arbitrary duration -- on HPC platforms, providing high throughput and high resource utilization. RAPTOR supports the high throughput virtual screening requirements of DOE's National Virtual Biotechnology Laboratory effort to find therapeutic solutions for COVID-19. RAPTOR has been used on $>8000$ compute nodes to sustain 144M/hour docking hits, and to screen $\sim$10$^{11}$ ligands. To the best of our knowledge, both the throughput rate and aggregated number of executed tasks are a factor of two greater than previously reported in literature. RAPTOR represents important progress towards improvement of computational drug discovery, in terms of size of libraries screened, and for the possibility of generating training data fast enough to serve the last generation of docking surrogate models.

Biotechnology2022arXiv
Periodicals

Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization

Protein language models have emerged as powerful tools for sequence generation, offering substantial advantages in functional optimization and denovo design. However, these models also present significant risks of generating harmful protein sequences, such as those that enhance viral transmissibility or evade immune responses. These concerns underscore critical biosafety and ethical challenges. To address these issues, we propose a Knowledge-guided Preference Optimization (KPO) framework that integrates prior knowledge via a Protein Safety Knowledge Graph. This framework utilizes an efficient graph pruning strategy to identify preferred sequences and employs reinforcement learning to minimize the risk of generating harmful proteins. Experimental results demonstrate that KPO effectively reduces the likelihood of producing hazardous sequences while maintaining high functionality, offering a robust safety assurance framework for applying generative models in biotechnology.

Biotechnology2025arXiv
Periodicals

Illuminating Protein Dynamics: A Review of Computational Methods for Studying Photoactive Proteins

Photoactive proteins absorb light and undergo structural changes that enable them to perform essential biological functions. These proteins are critical for understanding light-induced biological processes, making them important in biophysics, biotechnology, and medicine. One effective approach to uncovering photoactive processes is through computational methods. These techniques provide atomic-level insights into the structural, electronic, and dynamic changes that occur upon light absorption. By employing these methods, we can gain a better understanding of processes that are challenging to capture experimentally, such as chromophore isomerization and protein conformational changes. Here, we provide a brief overview of the different families of photoactive proteins and the computational methods used to study them, including bioinformatics, molecular dynamics, and enhanced sampling. Our review can serve as an introduction to computational methods for studying light-activated molecular processes, specifically targeting researchers beginning their journey in this field.

Biotechnology2025arXiv
Periodicals

FDTD Simulation of Exposure of Biological Material to Electromagnetic Nanopulses

Ultra-wideband (UWB) electromagnetic pulses of nanosecond duration, or nanopulses, are of considerable interest to the communications industry and are being explored for various applications in biotechnology and medicine. The propagation of a nanopulse through biological matter has been computed in the time domain using the finite difference-time domain method (FDTD). The approach required existing Cole-Cole model-based descriptions of dielectric properties of biological matter to be re-parametrized using the Debye model, but without loss of accuracy. The approach has been applied to several tissue types. Results show that the electromagnetic field inside a biological tissue depends on incident pulse rise time and width. Rise time dominates pulse behavior inside a tissue as conductivity increases. It has also been found that the amount of energy deposited by 20 $kV/m$ nanopulses is insufficient to change the temperature of the exposed material for the pulse repetition rates of 1 $MHz$ or less.

Biotechnology2004arXiv
Periodicals

MOLIERE: Automatic Biomedical Hypothesis Generation System

Hypothesis generation is becoming a crucial time-saving technique which allows biomedical researchers to quickly discover implicit connections between important concepts. Typically, these systems operate on domain-specific fractions of public medical data. MOLIERE, in contrast, utilizes information from over 24.5 million documents. At the heart of our approach lies a multi-modal and multi-relational network of biomedical objects extracted from several heterogeneous datasets from the National Center for Biotechnology Information (NCBI). These objects include but are not limited to scientific papers, keywords, genes, proteins, diseases, and diagnoses. We model hypotheses using Latent Dirichlet Allocation applied on abstracts found near shortest paths discovered within this network, and demonstrate the effectiveness of MOLIERE by performing hypothesis generation on historical data. Our network, implementation, and resulting data are all publicly available for the broad scientific community.

Biotechnology2017arXiv
Periodicals

The Future of Decoding Non-Standard Nucleotides: Leveraging Nanopore Sequencing for Expanded Genetic Codes

Expanding genetic codes from natural standard nucleotides to artificial non-standard nucleotides marks a significant advancement in synthetic biology, with profound implications for biotechnology and medicine. Decoding the biological information encoded in these non-standard nucleotides presents new challenges, as traditional sequencing technologies are unable to recognize or interpret novel base pairings. In this perspective, we explore the potential of nanopore sequencing, which is uniquely suited to decipher both standard and non-standard nucleotides by directly measuring the biophysical properties of nucleic acids. Nanopore technology offers real-time, long-read sequencing without the need for amplification or synthesis, making it particularly advantageous for expanded genetic systems like Artificially Expanded Genetic Information Systems (AEGIS). We discuss how the adaptability of nanopore sequencing and advancements in data processing can unlock the potential of these synthetic genomes and open new frontiers in understanding and utilizing expanded genetic codes.

Biotechnology2024arXiv
Periodicals

Sloppiness, robustness, and evolvability in systems biology

The functioning of many biochemical networks is often robust -- remarkably stable under changes in external conditions and internal reaction parameters. Much recent work on robustness and evolvability has focused on the structure of neutral spaces, in which system behavior remains invariant to mutations. Recently we have shown that the collective behavior of multiparameter models is most often 'sloppy': insensitive to changes except along a few 'stiff' combinations of parameters, with an enormous sloppy neutral subspace. Robustness is often assumed to be an emergent evolved property, but the sloppiness natural to biochemical networks offers an alternative non-adaptive explanation. Conversely, ideas developed to study evolvability in robust systems can be usefully extended to characterize sloppy systems.

Biotechnology2008arXiv
Periodicals

SwitchCraft: A Programmatic Framework for Designing State-Switching Proteins

Multistate mechanisms underlie many of the complex functions observed in natural proteins. The ability to rationally design multistate proteins would have transformative implications for many areas of biotechnology, yet lies beyond the capabilities of existing deep learning frameworks for protein design. To address this gap, we introduce SwitchCraft, a versatile and programmatic framework for designing state-switching proteins based on backpropagation through compositional design constraints parameterized by structure prediction models. In silico evaluations demonstrate success on a wide range of state-switching functional primitives, from allosteric regulation of motifs to discrimination of bound ligand identities. Using these primitives, we demonstrate an in silico strategy for de novo design of fluorescent biosensors to arbitrary small molecule analytes. These results position SwitchCraft at the inception of a powerful paradigm for higher-order functional protein design. Code is available at https://github.com/bjing2016/switchcraft.

Biotechnology2026arXiv
Periodicals

Guidelines for reporting the use of gel electrophoresis in proteomics

the MIAPE Gel Electrophoresis (MIAPE-GE) guidelines specify the minimum information that should be provided when reporting the use of n-dimensional gel electrophoresis in a proteomics experiment. Developed through a joint effort between the gel-based analysis working group of the Human Proteome Organisation's Proteomics Standards Initiative (HUPO-PSI; http://www.psidev.info/) and the wider proteomics community, they constitute one part of the overall Minimum Information about a Proteomics Experiment (MIAPE) documentation system published last August in Nature Biotechnology

Biotechnology2009arXiv
Periodicals

Predictive Entropy Search for Efficient Global Optimization of Black-box Functions

We propose a novel information-theoretic approach for Bayesian optimization called Predictive Entropy Search (PES). At each iteration, PES selects the next evaluation point that maximizes the expected information gained with respect to the global maximum. PES codifies this intractable acquisition function in terms of the expected reduction in the differential entropy of the predictive distribution. This reformulation allows PES to obtain approximations that are both more accurate and efficient than other alternatives such as Entropy Search (ES). Furthermore, PES can easily perform a fully Bayesian treatment of the model hyperparameters while ES cannot. We evaluate PES in both synthetic and real-world applications, including optimization problems in machine learning, finance, biotechnology, and robotics. We show that the increased accuracy of PES leads to significant gains in optimization performance.

Biotechnology2014arXiv
Periodicals

The Thing With E.coli: Highlighting Opportunities and Challenges of Integrating Bacteria in IoT and HCI

With advances in nano- and biotechnology, bacteria are receiving increasing attention in scientific research as a potential substrate for Internet of Bio-Nano Things (IoBNT), which involve networking and communication through nanoscale and biological entities. Harnessing the special features of bacteria, including an ability to become autonomous - helped by an embedded, natural propeller motor - the microbes show promising array of application in healthcare and environmental health. In this paper, we briefly outline significant features of bacteria that allow analogies between them and traditional computerized IoT device to be made. We argue that such comparisons are critical in terms of helping researchers to explore human-bacteria interaction in the context of IoT and HCI. Furthermore, we highlight the current lack of tangible infrastructure for researchers in IoT and HCI to access and experiment with bacteria. As a potential solution, we propose to utilize the DIY biology movement and gamification techniques to leverage user engagement and introduction to bacteria.

Biotechnology2019arXiv
Periodicals

Controllable protein design with particle-based Feynman-Kac steering

Proteins underpin most biological function, and the ability to design them with tailored structures and properties is central to advances in biotechnology. Diffusion-based generative models have emerged as powerful tools for protein design, but steering them toward proteins with specified properties remains challenging. The Feynman-Kac (FK) framework provides a principled way to guide diffusion models using user-defined rewards. In this paper, we enable FK-based steering of RFdiffusion through the development of guiding potentials that leverage ProteinMPNN and structural relaxation to guide the diffusion process towards desired properties. We show that steering can be used to consistently improve predicted interface energetics and increase binder designability by $89.5\%$. Together, these results establish that diffusion-based protein design can be effectively steered toward arbitrary, non-differentiable objectives, providing a model-independent framework for controllable protein generation.

Biotechnology2025arXiv
Periodicals

Dynamics of magnetic nano-flake vortices in Newtonian fluids

We study the rotational motion of nano-flake ferromagnetic discs suspended in a Newtonian fluid, as a potential material owing the vortex-like magnetic configuration. Using analytical expressions for hydrodynamic, magnetic and Brownian torques, the stochastic angular momentum equation is determined in the dilute limit conditions under applied magnetic field. Results are compared against experimental ones and excellent agreement is observed. We also estimate the uncertainty in the orientation of the discs due to the Brownian torque when an external magnetic field aligns them. Interestingly, this uncertainty is roughly proportional to the ratio of thermal energy of fluid to the magnetic energy stored in the discs. Our approach can be implemented in many practical applications including biotechnology and multi-functional fluidics.

Biotechnology2016arXiv
Periodicals

Molecular docking studies on Jensenone from eucalyptus essential oil as a potential inhibitor of COVID 19 corona virus infection

COVID-19, a member of corona virus family is spreading its tentacles across the world due to lack of drugs at present. However, the main viral proteinase (Mpro/3CLpro) has recently been regarded as a suitable target for drug design against SARS infection due to its vital role in polyproteins processing necessary for coronavirus reproduction. The present in silico study was designed to evaluate the effect of Jensenone, a essential oil component from eucalyptus oil, on Mpro by docking study. In the present study, molecular docking studies were conducted by using 1-click dock and swiss dock tools. Protein interaction mode was calculated by Protein Interactions Calculator.The calculated parameters such as binding energy, and binding site similarity indicated effective binding of Jensenone to COVID-19 proteinase. Active site prediction further validated the role of active site residues in ligand binding. PIC results indicated that, Mpro/ Jensenone complexes forms hydrophobic interactions, hydrogen bond interactions and strong ionic interactions. Therefore, Jensenone may represent potential treatment potential to act as COVID-19 Mpro inhibitor. However, further research is necessary to investigate their potential medicinal use.

Biotechnology2020arXiv
Periodicals

Emerging categories in scientific explanations

Clear and effective explanations are essential for human understanding and knowledge dissemination. The scope of scientific research aiming to understand the essence of explanations has recently expanded from the social sciences to machine learning and artificial intelligence. Explanations for machine learning decisions must be impactful and human-like, and there is a lack of large-scale datasets focusing on human-like and human-generated explanations. This work aims to provide such a dataset by: extracting sentences that indicate explanations from scientific literature among various sources in the biotechnology and biophysics topic domains (e.g. PubMed's PMC Open Access subset); providing a multi-class notation derived inductively from the data; evaluating annotator consensus on the emerging categories. The sentences are organized in an openly-available dataset, with two different classifications (6-class and 3-class category annotation), and the 3-class notation achieves a 0.667 Krippendorf Alpha value.

Biotechnology2025arXiv
Periodicals

Approximation Algorithms for Minimum PCR Primer Set Selection with Amplification Length and Uniqueness Constraints

A critical problem in the emerging high-throughput genotyping protocols is to minimize the number of polymerase chain reaction (PCR) primers required to amplify the single nucleotide polymorphism loci of interest. In this paper we study PCR primer set selection with amplification length and uniqueness constraints from both theoretical and practical perspectives. We give a greedy algorithm that achieves a logarithmic approximation factor for the problem of minimizing the number of primers subject to a given upperbound on the length of PCR amplification products. We also give, using randomized rounding, the first non-trivial approximation algorithm for a version of the problem that requires unique amplification of each amplification target. Empirical results on randomly generated testcases as well as testcases extracted from the from the National Center for Biotechnology Information's genomic databases show that our algorithms are highly scalable and produce better results compared to previous heuristics.

Biotechnology2004arXiv
Periodicals

Opportunities at the interface of network science and metabolic modelling

Metabolism plays a central role in cell physiology because it provides the molecular machinery for growth. At the genome-scale, metabolism is made up of thousands of reactions interacting with one another. Untangling this complexity is key to understand how cells respond to genetic, environmental, or therapeutic perturbations. Here we discuss the roles of two complementary strategies for the analysis of genome-scale metabolic models: Flux Balance Analysis (FBA) and network science. While FBA estimates metabolic flux on the basis of an optimisation principle, network approaches reveal emergent properties of the global metabolic connectivity. We highlight how the integration of both approaches promises to deliver insights on the structure and function of metabolic systems with wide-ranging implications in discovery science, precision medicine and industrial biotechnology.

Biotechnology2020arXiv
Periodicals

Plasmons in graphene: Recent progress and applications

Owing to its excellent electrical, mechanical, thermal and optical properties, graphene has attracted great interests since it was successfully exfoliated in 2004. Its two dimensional nature and superior properties meet the need of surface plasmons and greatly enrich the field of plasmonics. Recent progress and applications of graphene plasmonics will be reviewed, including the theoretical mechanisms, experimental observations, and meaningful applications. With relatively low loss, high confinement, flexible feature, and good tunability, graphene can be a promising plasmonic material alternative to the noble metals. Optics transformation, plasmonic metamaterials, light harvesting etc. are realized in graphene based devices, which are useful for applications in electronics, optics, energy storage, THz technology and so on. Moreover, the fine biocompatibility of graphene makes it a very well candidate for applications in biotechnology and medical science.

Biotechnology2013arXiv
Periodicals

Nano-Biotechnology: Structure and Dynamics of Nanoscale Biosystems

Nanoscale biosystems are widely used in numerous medical applications. The approaches for structure and function of the nanomachines that are available in the cell (natural nanomachines) are discussed. Molecular simulation studies have been extensively used to study the dynamics of many nanomachines including ribosome. Carbon Nanotubes (CNTs) serve as prototypes for biological channels such as Aquaporins (AQPs). Recently, extensive investigations have been performed on the transport of biological nanosystems through CNTs. The results are utilized as a guide in building a nanomachinary such as nanosyringe for a needle free drug delivery.

Biotechnology2010arXiv