I Let an AI Dream Its Own Training Data. It Didn't Stop Improving.
In this article
Bottom line: A research project called Dream-RSI, trending on Hacker News this month, trains an AI model by having it generate — essentially "dream up" — its own practice scenarios, learn from the outcomes, and then generate the next batch, recursively improving without any new real-world data.
It works, for a while, which is unsettling, because the loop it's running has a much older name in psychology: rumination.
The uncomfortable finding isn't that a machine can get smarter by training on its own imagination — it's that it can only keep getting smarter if something outside the loop keeps checking the dream against reality.
That's the exact rule that decides whether your late-night rehearsal of an argument sharpens you or just wears a groove into the same wrong story.
A client of mine — I'll call her Priya — once told me she'd had the same argument with her old boss four hundred times. Not out loud. In the shower, in traffic, at 2 a.m.
staring at the ceiling. Every rehearsal, she found sharper lines, a tighter case for why she'd been wronged, a better comeback for the thing he'd said in that meeting.
By the time she actually ran into him at a conference, she was armed and ready.
And she completely misread the room, because all four hundred rehearsals were built on the same faulty memory, replayed and polished until it felt like fact. She hadn't practiced the conversation.
She'd trained herself on a hallucination.
I thought about Priya for the first time in years this week, reading about a machine learning technique making the rounds called Dream-RSI — recursive self-improvement through evolving worlds.
The premise: instead of feeding a model more human-labeled data, you let it generate its own training scenarios, learn from how those scenarios play out, then generate the next, harder round based on what it just learned.
No new external input. Just the system, improving on its own output, over and over.
The Loop We're Already Running
Here's the part that should land differently depending on whether you read it as a machine learning story or a Tuesday-night-in-your-head story: humans have been running some version of this loop for as long as we've had inner monologues.
We call it worrying. We call it planning.
We call it "just thinking it through." Psychologist Susan Nolen-Hoeksema spent decades studying what happens when self-generated replay goes unchecked, and her research on rumination found it doesn't just fail to solve problems — it actively worsens mood, narrows thinking, and predicts longer depressive episodes.
Meanwhile Pennebaker's expressive writing studies found the opposite effect from a structurally similar habit: people who wrote about a hard experience for just 15 minutes a day, four days running, showed measurably better immune function and fewer doctor visits in the following months.
Same basic move — generate your own data about your own life, learn from it, do it again. Wildly different outcomes.
If you've ever left a journaling habit or a "let me just think this through one more time" session feeling worse than when you started, you already know this loop can go either way.
The question nobody in the self-help aisle quite answers is what actually determines which way it goes.
What the Machines Accidentally Proved
Here's the reframe, and it's the part that made me sit up: machine learning researchers have known for a while that training a model recursively on its own synthetic output, with no fresh real-world data mixed back in, tends to produce something called model collapse — the outputs get narrower, more confident, and further from reality with every generation, even as the model's own sense of its accuracy goes up.
That is Priya's four hundred rehearsals, described in a systems paper.
The Dream-RSI approach gets around this specifically by building in reality anchors — checkpoints where the "dreamed" scenario gets scored against something outside the loop, some ground truth the model didn't generate itself.
Without that anchor, recursive self-improvement isn't improvement. It's just erosion that feels like progress, because the thing measuring the progress is the same thing that made it up.
This is worth sitting with, because it flips the standard self-help advice on its head.
The problem was never that Priya rehearsed. Mental rehearsal is a documented, legitimate tool — surgeons and Olympic athletes use structured visualization alongside real practice and measurably outperform those who only do one or the other.
The problem was that her rehearsal never touched anything outside her own head.
No fresh information. No outside check.
Just the same story, dreaming itself smarter, growing more convinced and less accurate with every pass — which is exactly what happens when people use a chatbot as a sounding board and it just agrees with whatever they already believe, a pattern that's become common enough to have its own name now.
The DREAM Loop
If self-generated growth only works with a reality anchor, then the fix isn't "stop rehearsing" — it's building the anchor in on purpose.
I've started calling this the DREAM Loop with clients, partly because it's memorable and partly because the acronym is doing real work.
D — Dream deliberately
Pick one specific situation you keep replaying or preparing for — a conversation, a decision, a pitch. Give yourself an actual window (10 minutes, not an open-ended spiral) to run it mentally.
Deliberate rehearsal with a time box behaves completely differently than rumination with no edges.
R — Reality-check it
Before you trust anything the rehearsal produced, hold it up against something outside your own head. Ask the actual person a clarifying question. Check the number.
Read the email again instead of the version of it you remember. This is the single step Priya skipped four hundred times.
E — Extract the signal
Most rehearsals generate 90% noise and one real insight.
Write down the one thing that changed — a question you should actually ask, a fact you were missing, a fear that turned out to be the whole thing.
Discard the rest instead of carrying it into the next round.
A — Anchor it in action
Translate the distilled insight into one small, real, external action within 24 hours. Send the message. Ask the question.
Make the call. This is what stops the loop from folding back into itself — it forces new, real data into the next round instead of just more imagination.
M — Measure and repeat
Notice what actually happened when the real world responded, and let that — not your prediction of it — be the input for the next dream.
This is the step that turns rehearsal into something closer to training, instead of a closed loop dreaming itself into a corner.
What This Looks Like on a Tuesday
Try this the next time you catch yourself rehearsing something for the third or fourth time. Set a 10-minute timer and actually run the scenario — out loud, or written, doesn't matter.
When the timer ends, force one reality-check question: what's one thing here I'm assuming instead of knowing? Then go find that one piece of real information before you rehearse again.
If you're prepping for something bigger — a hard conversation, a negotiation, a performance review — do exactly one dream cycle, then take one small real action (a text, a five-minute call, a single question to someone who'd actually know), and only then allow yourself a second rehearsal, now built on real information instead of your first guess dressed up as memory.
This is also, incidentally, the difference between people who get sharper from talking things through with an AI chat assistant like Claude 4.6 or ChatGPT 5 and people who get more confidently wrong.
The tool isn't the variable. The reality anchor is.
The Loop You're Already In
Priya eventually got the conversation right — not because she rehearsed it a four-hundred-and-first time, but because she finally asked a mutual colleague what had actually happened in that meeting, and the story she'd been training on for months turned out to be half true at best.
One real data point undid four hundred imagined ones.
The machines just proved this at scale, with cleaner math than any of us get in a shower at 7 a.m.
Where in your life are you running the four-hundredth rehearsal on the same faulty memory, mistaking confidence for accuracy because nothing outside your head has checked it in a while?


