Expression systems built around your therapy

Lap-in-the-loop training data

Predictive models for gene expression are only as good as the data behind them. Annogen generates large-scale, functionally measured sequence-to-expression datasets in the exact cell types, vector contexts and conditions your programs run in, so your machine-learning models learn from data that reflects your biology rather than public sequences borrowed from other systems.

What it involves

Public regulatory-genomics data is broad but shallow for any specific therapeutic context. It rarely covers your cell type, your vector or the expression behavior you actually need to predict, so a model trained on it generalizes poorly to the decisions that matter: which non-coding variant shifts expression, or which synthetic promoter will be both strong and specific in a primary T cell.

Annogen closes that gap by measurement. Using SuRE™, we quantify the activity of hundreds of millions of regulatory sequences in a single experiment, directly in the relevant cell type and vector configuration. Each element is tracked by dozens to hundreds of barcodes, so every label in the dataset is an average over many independent measurements. That redundancy is what makes the data suitable for training: low label noise, a wide dynamic range, and reference and control sequences built in, so values are calibrated rather than only relative.

The result is a dataset built around your biology. Sequence in, measured expression out, in the system you care about, with the density and quality that sequence-to-function models need.

Lab in the loop

Lab in the loop means the wet lab is part of the training cycle, not a one-off data drop. Your model proposes sequences, we synthesize and measure them with SuRE™, and the measured results feed the next round of training and design. Annogen already runs this loop internally: we use our own screen data to design fully synthetic promoters, then validate them experimentally before anything is called a hit. We can run the same closed loop around your models and your targets, so each cycle sharpens both the model and the candidate set.

Annogen closes that gap by measurement. Using SuRE™, we quantify the activity of hundreds of millions of regulatory sequences in a single experiment, directly in the relevant cell type and vector configuration. Each element is tracked by dozens to hundreds of barcodes, so every label in the dataset is an average over many independent measurements. That redundancy is what makes the data suitable for training: low label noise, a wide dynamic range, and reference and control sequences built in, so values are calibrated rather than only relative.

The result is a dataset built around your biology. Sequence in, measured expression out, in the system you care about, with the density and quality that sequence-to-function models need.

What you can train

Variant-effect prediction
Regulatory element and sequence design
Cell-type specificity models
Expression optimization for manufacturing
Promoter and enhancer libraries
Saturation mutagenesis datasets
Custom dataset generation

Dataset types

Why Annogen

Measured, not predicted

Every label is an experimental readout of activity in living cells, not a model output.

Built for your biology

Your cell type, your vector, your conditions, rather than a generic public set.

Built for training

Barcode redundancy gives high signal-to-noise and calibrated, low-noise labels, which matters more for model quality than raw row count.

Yours to keep

The dataset, and any model you train on it, are your intellectual property. Annogen retains the SuRE™ platform, not your data. This is a deliberate part of how we work, and a common reason clients choose us for model training.

How we work

Define the prediction problem and the biology: what your model needs to predict, in which cell type, vector and conditions.

Design the library to cover the sequence space your model needs to learn, with references and controls spiked in for calibration.

Measure at scale with SuRE™, each element tracked by many barcodes.

Deliver the dataset with the underlying design and quality-control metadata, ready for training.

Close the loop where wanted, measuring model-proposed sequences in successive rounds.

Interested?

Let’s start with your ambitions. Tell us what your models need to predict and in which system. We will help you explore how measured, lab-in-the-loop data can sharpen them.

Is this real measured data or predictions?
How large are the datasets?
Do we own the data and the models?
Can you match our existing data format?

Frequently Asked Questions

We imagine there could be some questions you want to ask us. Discover the most frequently asked questions about this subject right here. 

This website uses cookies

We use cookies to personalize content, provide social media features, and analyze our traffic. We also share information about your use of our site with our analytics partners. You can change your preferences at any time. For more information, please see our Privacy Policy and Cookie Policy.