Science as a Jungle


All known terrestrial life uniformly uses unique strings of biological sequence as their mechanism of inheritance. When born, each organism is encoded by some set of nucleic acid strings of just four to five chemicals {A, C, G, T, U}. Despite this small alphabet, combinations of these characters completely encode all organisms/lifestyles present on planet Earth. Independent of this fact, large-language models, solely trained on the unique set of sequences ({tokens}) produced by humans from prehistory to now, are able to: solve Millennium problems, cooperate in swarms to hack servers, and (maybe) yield significant progress and discovery. By applying the same tools that yielded this progress in LLMs to now biological tokens would mean creating a model capable of reading, understanding, and then generating novel sequence. Unfortunately, our current sampling of this sequence space is extremely biased. Public sequencing data is still in 'prehistory'. It is overwhelmingly dominated by a couple organisms. Human genomes, yeast genomes, pathogen genomes. Organisms that humans are predisposed to spend money to sequence, characterize, and study. Free from compute and lab constraints, I would propose leading a systematic sequencing and characterization workflow to bring biological knowledge into the modern era. I would begin by identifying the current gaps in locations/life where humans have neither read nor experimented on biological sequences. These gaps would then be filled with a combination of unbiased environmental metagenomic sequencing and partnering with CROs to functionally characterize biological sequences in a lab. The goal would be to build a map between a biological sequence to both understanding and generating functional, novel sequence. It would be of genuine interest to characterize the biological equivalent of scaling laws. LLM scaling laws demonstrate how model performance relates to training data, compute, and matrix size. How does biological language model (bLM) performance scale with newly sampled environmental sequences, deep lab characterization, and parameter size? from compute and lab constraints, The most ambitious project I would pursue at Anthropic would be a attempt to map the limits of our ability to predict novel function from sequence alone. LLM scaling laws map large-language model performance to training data, compute, and parameter size; I want to identify the biological equivalent of these same relationships. Current biological foundation models remain hamstrung by the severe sampling bias present in public databases, an open question is quantifying how much new signal can be learned by expanding these databases to non-model organisms and more holistic experimentation. Working with experimental biologists, I would map biological sequences to experimentally measured functions and then compare model performance as we independently expand sequence coverage, functional evidence, and model capacity. I would test whether functional discovery and predictive performance saturate together, and whether new experimental observations allow further improvement. For Anthropic, this would help distinguish limits of current models from limits of what we have measured, and determine when progress requires more compute, broader sequencing, or different experiments.



___________________________________________________________________

Arya Kaul (C) now - forever -> more essays