Breakthrough

US Researchers Map 22 Million Immune Cells at Genome Scale in Primary Human T Cells, Releasing Largest Functional Genomics Dataset for AI Model Training

A genome-scale perturbation study knocking out 12,800 genes one by one across primary human T cells has produced the largest functional immune cell dataset to date, freely released to accelerate AI virtual cell model development and immune circuit discovery globally.

US Researchers Map 22 Million Immune Cells at Genome Scale in Primary Human T Cells, Releasing Largest Functional Genomics Dataset for AI Model Training

InnoDexis has published its latest Innovation Intelligence Report covering functional genomics and immune cell biology, analyzing a landmark dataset innovation from the United States. The report reveals that researchers at Gladstone-UCSF, Stanford, and UCSF have produced the largest genome-scale perturbation dataset in primary human immune cells to date — knocking out nearly 12,800 genes individually across primary human T cells to generate 22 million high-quality analysed cells. The dataset has been released globally as a functional lookup table to support AI virtual cell model training and immune gene circuit discovery, representing the largest single contribution to Biohub's Billion Cells Project.

Key Findings

Nearly 12,800 genes were knocked out one by one across primary human T cells sourced from real blood donors, generating 33.4 million total cells screened and 22 million high-quality cells retained for analysis. This scale of genome-wide perturbation in primary human cells — rather than artificial cell lines — had not previously been achieved, establishing this dataset as the first of its kind at genome scale in a primary human immune cell context.

The dataset captures gene circuit behaviour across both resting and infection-fighting cell states, preserving the dynamic nature of immune cell function. Previous genomics efforts catalogued gene expression in static conditions — this study measures what occurs when each gene is switched off across distinct immune states, shifting the analytical frame from observational description to causal functional mapping.

Primary human T cells were used throughout, preserving real patient immune variation rather than the artificial uniformity of cell lines. This design choice makes the findings directly comparable to disease states encountered in clinical settings, increasing the dataset's relevance for translational research in cancer immunotherapy and autoimmune disease relative to prior efforts conducted in simplified model systems.

The complete dataset has been released globally and freely, functioning as a functional lookup table for the broader scientific field. Its release is explicitly designed to support AI virtual cell model training — positioning it as infrastructure for the next generation of computational biology tools rather than a self-contained experimental result. It represents the largest single contribution to date to Biohub's Billion Cells Project.

The dataset enables discovery of gene circuits governing immune cell behaviour at a resolution and scale not previously available. Where earlier genomics work described which genes are expressed, this perturbation approach measures what each gene actually does — a distinction that directly accelerates target identification for therapeutic programmes in immunology and oncology.

Strategic Insight and Trend Analysis

The dataset represents a structural transition point in the life sciences — what the researchers characterise as a third wave of genomics, moving from sequencing and cataloguing toward understanding what genetic changes actually do. The first wave produced the human genome sequence. The second wave produced large-scale catalogues of gene expression across tissues and cell types. This dataset inaugurates a third phase: causal functional mapping at genome scale in clinically relevant primary human cells.

The strategic weight of this transition lies in its implications for AI-driven drug discovery. Machine learning models trained on observational gene expression data are constrained by the correlational nature of that data — they can identify patterns but cannot reliably distinguish causal relationships from associations. A genome-scale perturbation dataset in primary human T cells provides a fundamentally different training signal: functional, causal, and generated under conditions that preserve real patient variation. This distinction matters enormously for the predictive validity of AI virtual cell models.

The decision to release the dataset freely and globally compounds its strategic impact. Rather than functioning as a proprietary asset for a single institution or company, it operates as shared field infrastructure — accelerating progress across laboratories simultaneously. For AI model development specifically, the dataset's scale of 22 million high-quality cells from 12,800 gene knockouts creates a training corpus of sufficient depth to support meaningful generalisation across immune cell biology.

The implications for target discovery timelines in cancer immunotherapy and autoimmune research are significant. If AI models trained on this dataset can reliably predict the functional consequences of gene perturbations across immune states, the iterative cycle of hypothesis, experiment, and validation that currently defines early-stage drug discovery could compress materially.

Global and Industry Implications

For corporates and R&D teams in pharmaceutical and biotechnology research, the freely released dataset represents an immediately accessible resource for immune target discovery programmes. Teams working on cancer immunotherapy, autoimmune therapies, and immune-oncology combination strategies can integrate this functional map into computational target prioritisation workflows without requiring proprietary data generation at equivalent scale.

For investors and capital allocators, the dataset signals accelerating maturation of AI-driven functional genomics as a drug discovery platform. Organisations building AI virtual cell models or perturbation-based target discovery platforms now have access to a genome-scale primary human immune cell training corpus — a resource that meaningfully raises the ceiling for what such models can achieve and compresses the data acquisition timelines that previously constrained this approach.

For policymakers and national innovation bodies, the multi-institution collaboration between Gladstone-UCSF, Stanford, and UCSF — supported through Biohub's Billion Cells Project — demonstrates the capacity of coordinated public research infrastructure investment to produce shared scientific assets with global utility. The open-release model adopted here provides a template for maximising the downstream impact of large-scale genomics investments.

InnoDexis Statement

"The genome-scale functional map of primary human immune cells marks a structural shift in genomics — from cataloguing gene expression to measuring causal gene function at clinical scale, creating shared infrastructure that could accelerate AI-driven immune target discovery across the field," noted InnoDexis in its latest intelligence report.

Conclusion

As AI virtual cell models mature and functional genomics datasets grow in scale and clinical relevance, the availability of a genome-wide perturbation map in primary human T cells establishes a new reference point for the field. The transition from observational cataloguing to causal functional mapping at genome scale will shape how target discovery, immune circuit analysis, and therapeutic hypothesis generation are conducted across cancer immunotherapy and autoimmune research in the years ahead. InnoDexis will continue to monitor developments in functional genomics, AI virtual cell platforms, and perturbation-based drug discovery. The complete Functional Genomics Innovation Intelligence Report is available to InnoDexis subscribers and enterprise clients.

About InnoDexis

InnoDexis is a global Innovation Intelligence platform that tracks, analyzes, and interprets breakthrough innovations, prototypes, and emerging technologies across industries and countries. Its intelligence helps corporates, investors, and policymakers understand the true structure and direction of global innovation. Learn more at innodexis.ai.

Ready to go beyond this brief?