Can a model learn directly from the world as it is?
Learning from dirty, incomplete human data without an expensive cleaning pipeline.
Introducing TasteMaxxer 1
TasteMaxxer 1 is a large language model that learns from messy human data—and develops a taste of its own.

The work ahead
Learning from dirty, incomplete human data without an expensive cleaning pipeline.
Continual learning through weight updates, without catastrophic forgetting.
Learning through reinforcement without millions of environment rollouts.
Join the early list for notes from the lab and news on TasteMaxxer 1.