Room 3 of 5

Training

Nobody wrote rules for every word in the English language. Instead, they did something faster and crazier.

They fed it the internet. Billions of pages. News. Medical journals. Legal briefs. Reddit threads. Romance novels. Earnings calls. Everything ever written and posted and forgotten.

Then they refined it further — human reviewers ranked its outputs, teaching it which responses people actually prefer. This is why modern AI sounds helpful rather than chaotic.

The toggles below represent categories of training data. They’re all on right now. Start turning them off.

In reality, training takes weeks and billions of dollars. These toggles simulate the effect — the principle is real, even if the speed is not.

Turn off everything except “Legal filings.” Read the response. The model sounds like a lawyer. Not a bad lawyer. A convincing lawyer. It uses words like “pursuant” and “fiduciary obligations.”

Now turn off everything except “Social media.” The model sounds like your nephew’s Twitter feed. Emoji. Slang. Opinions delivered with the confidence of someone who has never been wrong because they’ve never been specific.

The model became its training data. Completely. When you changed the training data, you changed the model’s entire personality.

When someone tells you their model is “unbiased,” ask one question: What was it trained on? If they won’t tell you, the answer is “the entire internet, including the worst parts of it.”

What this means when you use AI

Every model has a knowledge cutoff — a date after which it knows nothing. If you’re asking about recent events, you need to supply the information yourself.

When a model gives you investment advice, legal analysis, or medical guidance, ask: was this in its training data? If you’re working with proprietary information the model has never seen, tell it explicitly. Don’t assume it knows what you know.

A “general purpose AI” and an “AI for finance” are different products, not different marketing. The training data is the product.

Room 4: Scale

So training data shapes the model. What happens when you make it bigger?

Continue to Scale