Posts

Modelling genomic language using NLP and LLMs

Image
Author : Navya Tyagi  Genomic data consists of DNA, RNA, and protein sequences that can be represented as strings of unstructured text. These sequences can be very large in size. For instance, human DNA is made of 3 billion A,G,C, and T letters. There are hidden patterns that can be considered equivalent to "words" in a natural language. But all of these words are not known and more importantly the grammar that genomic lanaguege follow is not well understood. These biological words with critical functions are of interest to study disease and development processes. Sometime a mutation in these "words" may result in a disease condition. Determining these "words" with biological function is a computational challenge. Figure 1: a) Code of DNA can be written using letters A,C,G,T. A pairs wiht T and G pairs with C making a double stnraded (helical) structure out of two DNA strands. b) message from DNA is transcribed to RNA (also known as messenger RNA). The ...

Clinical data standards for large-scale linking of Big biomedical and health data

Image
Data is the basic unit of information and in case of health and clinical data this comes in many shapes and sizes. The clinical data varies from numerical measurements and quantities, to image, documentation, digital codes and narrative text with facts and observations. This data becomes available to a learning or accountability system by various means. For example,  it can be data acquisition from paper records, direct entry into a computer system, or reuse of data collected by others. The other aspect is the high throughput data generated from R&D. Thus, we can safely say that the data is Big, multi dimensional and multimodal. The individual data elements can be grouped based on a common criteria to form datasets e.g. vital measurement data from patients electronic health records. The data from biological and health domain is growing at the fastest rate being >90% of the data generated each year. It needs no convincing that big data will increase efficiency and accounta...

Common misconceptions and pitfalls of using ROC

Image
  Receiver operating curve (ROC)  An ROC curve is a commonly used technique to visualise, organise, and select or compare classifiers based on their performance. Where a classifier is usually binary with two possible outcomes. Historically, the use of ROCs comes from WW-II where it was used for assessing performance of signal detection and specifically to find out whether a signal from radar was a true positive or a false positive. The guy who operated the radar receiver would be known as ‘receiver operator’, and hence, ROC got its name :) Sometime in the 1970s ROC started making its appearance in the field of medicine, where it was used to evaluate and compare algorithms. Now, I will elaborate on understanding different components of an ROC, what they mean, and how they are used. Let's take an example  where an observed instance is mapped to one of two class labels, say COVID positive or negative. This test can be performed based on a variety of factors a.k.a features. A...