- Mateo's Rabbit Burrow
- Posts
- The immune system doesn't need a better weapon. It needs cancer's address.
The immune system doesn't need a better weapon. It needs cancer's address.
Previously, I asked whether cancer is a navigational problem or a destruction problem.
Is the cell lost, or too far gone to come back?
I ended it with inversion. Don't chase the moment cancer arrives. Map health. Map full malignancy. Subtract both. Whatever survives is the stage that matters.
So I went looking for that in pancreatic cancer, calling it Project Calchas.
And found something I wasn't looking for.
Right now, somewhere in you, a cell has made a mistake. A T cell will find it and kill it. Today, without you noticing.
Your immune system has been killing cancer your entire life.
So when it fails, the question isn't how do we kill harder. It's: what did the immune system stop seeing?
Every cell keeps a shop window
A cell chops up the proteins it makes inside and holds the pieces on its surface. Like a shop putting items in the window. Whatever it's building, it shows outside.
The window has a name: the HLA. Everyone's is slightly different, which will matter in a minute.
T cells walk the street and read windows.
A tumor cell has mutations, so it builds proteins with errors in their sequence. Those fragments reach the window the same way, except no healthy cell would ever display them.
That fragment is a neoantigen.
A flag that says: something in here is not me.
A T cell reads the flag and kills the cell.
That's the entire promise of cancer vaccines and T cell therapy. Not a stronger poison. A killer that already exists, handed the right address.
Then pick a handful of neoantigens.
A tumor carries thousands of mutations. A vaccine can carry a few dozen. A cell therapy carries one or two.
Everything collapses into one decision: which ones?
Every tool in the field answers the same way. Measure how tightly the fragment sits in the window. Binding. Clean, measurable, and the models for it are great.
But sitting in the window doesn't make a sale (the T cell still has to respond).
On the most realistic benchmark we have, fewer than 1 in 10 tested candidates turn out to be immunogenic.
Teams spend months and real money on targets the immune system quietly walks past. The number they trusted was a binding score wearing the costume of confidence.
That gap has a name. Neoantigen immunogenicity prediction. Every immunotherapy coming down the pipe is standing on it.
Munger's point was never to swim harder. It's to pick the right wave.
The chain
The field has known for years that binding isn't the answer. So it went to work. Thousands of people. Serious money. Better models.
Nothing beat the binding score it was built to improve on.
When that many people stop in the same place, it stops being a method problem.
Here's why. A T cell response isn't one event. It's a chain, and every link has to hold.
The mutated gene is switched on (we see this)
Enough tumor cells carry the mutation (we see this)
The cell chops the protein and ships the pieces to the window (modelled)
The fragment sits in the window and stays (modelled)
Enough copies reach the surface (blind)
This patient owns a T cell that can read it (blind)
That T cell survived thymic selection (partly)
The grip lasts long enough to pull the trigger (blind)
The T cell isn't exhausted and the surroundings allow it (blind)
Four of nine links are invisible to the input.
Link 6 decides everything. If this patient owns no T cell that can read the flag, the chance of a response is zero, no matter how beautifully the fragment binds.
Link 7 is the one people get backwards, including me at first. T cells aren't trained on tumor mutations. They're trained on you. Any receptor reacting too strongly to your own proteins gets deleted in the thymus, years before the cancer exists. So a neoantigen that looks too much like self probably lost its reader long ago.
Every tool in the field, mine included, predicts a nine-step chain from the first four steps.
Why more compute doesn't fix it
Information theory has a rule: processing can't create information. If the answer isn't in the input, no model recovers it.
So the ceiling belongs to the input, not the architecture. That's a different wall than "biology is complicated." It's closer to conservation of energy.
And it leaves a fingerprint. When unrelated methods meet at the same number, that's not coincidence. It's a measurement.
I tried to break it four times, writing the kill criteria before each run. A learned ranker. Twenty times more training data. A cancer-only slice of it. Extra biochemical signals.
Four attempts, four negatives. All four switched off.
Meanwhile, a published deep learning model built for exactly this scores 0.500 on the realistic benchmark. A coin flip.
That's what an exhausted input looks like from the outside.

What I actually built (Epitax Bio)
A deterministic scoring function on top of a machine-learned presentation predictor, with the probability calibrated separately.
Meaning: the order of the shortlist is arithmetic you could redo on paper. Four constants, chosen by hand, written down. The presentation estimate underneath is the field's standard neural network.
On top of binding, it uses what the tumor tells you. Is the gene switched on? Do enough tumor cells carry the mutation? How far the fragment has drifted from the healthy original?
And if a patient's HLA sits too far from anything the tool was tested on, it issues no probability at all. It warns you before you spend the run, not after.
Honest about what it isn't: as a pure predictor, Epitax ties the field's best signal. It doesn't beat it.
I don't count that as failure.
Picture a lab that can afford twenty tests.
Choose your ranking tool by the score everyone publishes, and those twenty tests turn up three real targets. Choose one that scores worse, and the same twenty turn up five maybes.
The new score that wins papers isn't the score that saves money.
Accuracy isn't the product. Trust is.
A refused answer is worth more than a confident wrong one.
From here, the plan is boring on purpose. Validate on data I never trained on, including a third benchmark built by another lab. Publish, including the results that come back empty. Put it in front of the people who pick these targets for living and ask where it breaks. Then grant applications.
Every step buys the same thing: permission to ask the real question.
Not "is this neoantigen immunogenic in general." That's close to unanswerable from the neoantigen and HLA window alone.
But "can this patient's T cells see it?"
Nobody holding only tumor data can ask that. And for the first time, sequencing a patient's receptors is cheap enough.
That's the problem Epitax Bio is after.
My goal hasn't moved.
Epitax is more than a step toward it, but I hold it that way on purpose. Fast. Pointed at one problem. Giving away as much as I can while I learn the ground under my feet.
A "practice project," with real stakes.
Check it out: epitax.bio (research use only).
If you pick neoantigens for a living, where the order breaks from your intuition.
Take care)
Mateo