---
title: A Tenth of Your Proteins Is Enough for the Immune System to Learn Not to Attack You
description: CSHL researchers show the immune system needs to sample only about 10% of the body's proteins to delete 90% of the T cells that could attack it.
author: Darie Nani (Editor-in-Chief)
updated: 2026-08-20T13:19:25.589Z
canonical: https://www.sovereignmagazine.com/article/immune-system-tolerance-sparse-sampling-generalization
image: https://cdn.nanimediahouse.com/cshl-immunoai-single-cell-map-181693.webp
categories: Science &amp; Tech
content_type: News
region: Global
publication: Sovereign Magazine
schema_type: Article
---

Your immune system faces a sorting problem with almost no margin for error. As T cells mature in the thymus, each one is tested against fragments of the body's own proteins, the self-peptides, and any T cell that binds too strongly to one is deleted. That process, negative selection, is how the immune system learns not to open fire on healthy tissue. The catch is that a single T cell only ever encounters a tiny fraction of the body's self-peptides. How it learns to tolerate the vast majority it never meets has been the hard part to explain.

Researchers at Cold Spring Harbor Laboratory report that the immune system clears this bar through generalization, and they have measured how little sampling it takes. Publishing in the peer-reviewed journal Science Advances, Hannah Meyer, who studies the thymus and T-cell development, and Saket Navlakha, who studies machine learning and biological computation, estimate that sparse, random sampling of only about 10% of the body's self-peptides is enough to correctly delete about 90% of self-reactive T cells.

## Sampling a Tenth of Self-Proteins Deletes Nine in Ten Dangerous T Cells

Using single-cell sequencing together with computational simulations, the team estimates that during its training a developing T cell interacts with only about 240 antigen-presenting cells drawn from a random sample of 2,000. That thin slice of the body's protein catalog is mathematically sufficient to remove roughly 90% of the T cells that would otherwise attack healthy tissue.

The explanatory frame the researchers use is borrowed openly from artificial intelligence. A machine-learning model trained on some photographs of dogs can still recognize dogs it was never shown, and the argument is that maturing T cells do something similar with self-peptides.

"Negative selection is a crucial process, but if T cells had to test against every single one of the body's peptides, it would take forever," Meyer said. "So, how do they learn to avoid friendly fire? We think it's through a process called generalization."

Generalization is not automatic. "Machine learning deals with the generalization problem all the time," Navlakha said, adding that it "is possible, but only under certain conditions. It's not black magic."

## Two Biological Conditions Make the Shortcut Work

The study identifies the conditions the immune system meets to make that shortcut reliable. First, the abundance of each self-peptide in the thymus closely mirrors its abundance in tissues around the body, which in machine-learning terms means the training data resemble the data the system will later be tested against. Second, T-cell receptors are cross-reactive: one receptor can recognize several similar peptides, so a T cell effectively learns about peptides it never directly encountered.

Reading negative selection as a learning algorithm is not itself a new proposal. Wortel and colleagues published "Is T Cell Negative Selection a Learning Algorithm?" in the journal Cells in 2020, and [applying machine-learning theory to immunology has been an active area for several years](https://www.frontiersin.org/journals/immunology/articles/10.3389/fimmu.2025.1651533/full). What the CSHL work adds is the measurement, the sparse-sampling numbers, and a test against disease.

## Modeling a Breakdown in the Process Reproduced a Real Autoimmune Disease

To ask what happens when generalization fails, the team turned to autoimmunity, the case where the immune system wrongly attacks healthy tissue. When they modeled failures of the process, the result reproduced features of autoimmune polyendocrine syndrome type 1 (APS-1), a rare autoimmune disease.

The link from a model to a clinic is not short. In animal studies, [impaired thymic negative selection does not always produce autoimmune disease](https://www.nature.com/articles/s41392-024-01952-8), and biological immune systems do things standard machine-learning models do not, including responding to danger and context signals rather than isolated antigens.

Navlakha frames the study as the start of a research program rather than a finished account. "We're calling this direction ImmunoAI," he said. "We're not trying to create AI inspired by the immune system, but we're studying how the immune system solves fundamental machine learning problems. When we start to look at the immune system as if it's another kind of AI, we may find out some surprising things about human health and disease."

Cold Spring Harbor Laboratory, a nonprofit biomedical research institution on Long Island, New York, was founded in 1890.

More on the research is at [Cold Spring Harbor Laboratory](https://www.cshl.edu/does-your-immune-system-learn-like-ai/).

## FAQ

**Q: What happens during negative selection of T cells?**
As T cells develop in the thymus, they are tested against fragments of the body's own proteins. Any T cell that binds too strongly to one of these self-peptides is deleted, which trains the maturing immune system not to attack healthy tissue.

**Q: What is autoimmune polyendocrine syndrome type 1 (APS-1)?**
APS-1 is a rare autoimmune disease, one in which the immune system attacks the body's own tissues. The team's model of a breakdown in central tolerance reproduced features of it.

**Q: Is applying machine learning to the immune system a new idea?**
No. Viewing T-cell negative selection as a learning algorithm was proposed in the peer-reviewed literature years ago, including a 2020 paper in the journal Cells, and computational immunology has been active for several years. The CSHL study's new contribution is the quantitative demonstration and the disease modeling, not the original concept.
