Opus 5.5 Just Quietly Automated the Job of AI Researchers

Bottom line: Anthropic's Opus 5.5, released this month, can now run a full research loop — read the literature, design an experiment, write and execute the code, and draft the findings — with only light human steering.

In a two-week test replicating three infrastructure research tasks I'd normally hand to a junior engineer, Opus 5.5 finished the literature-synthesis and experiment-design stages in under an hour each, work that used to eat two to three days.

It didn't replace judgment or novel hypothesis generation, and it still needed a human to catch two subtly wrong conclusions.

If your job involves reading papers, running benchmarks, and writing up results, that middle chunk of work just got a lot smaller.

I gave Claude Opus 5.5 a research question I already knew the answer to. Just to see if it would cheat, guess, or actually work through it.

It didn't cheat.

It read eleven papers, flagged which two contradicted each other, designed an experiment to resolve the contradiction, wrote the code to run it, and got an answer within about four percentage points of the number I already had sitting in a spreadsheet from six months of actual lab work.

That took ninety minutes.

I've been skeptical of "AI does research now" claims for two years. I'm less skeptical now.

The Setup: I Gave It My Actual Backlog

I run infrastructure for a mid-size platform team, and part of my job is a slow trickle of research-adjacent work: does this caching strategy actually reduce p99 latency under real traffic patterns, is this new retry backoff algorithm from a paper worth adopting, does batching inference requests this way trade off accuracy for throughput in a way we can live with.

It's not glamorous. It's reading, prototyping, measuring, writing a doc nobody reads carefully enough.

I picked three of these tasks sitting in my backlog as of September 2026 and ran each one twice — once the way I normally would, once handing the whole loop to Claude Opus 5.5 with minimal hand-holding.

I used the same prompt structure each time: define the question, point it at a folder of PDFs and our internal benchmarking harness, ask it to propose an experiment, then let it run.

The three tasks: evaluating a request-coalescing strategy for our inference gateway, comparing two approaches to KV-cache eviction under memory pressure, and checking whether a published speculative-decoding technique actually held up on our production model sizes rather than the toy models in the paper.

None of these are "solved problems" with a known right answer I could just look up.

They required reading, synthesizing, disagreeing with a paper's own conclusions in one case, and running real code against real data.

Article illustration

What Opus 5.5 Actually Does Differently

I've used earlier Claude models for research-assistant work before — summarizing papers, drafting literature reviews, that kind of thing. What's different with Opus 5.5 isn't any single capability.

It's that the loop closes without me babysitting each step.

It Chains Steps Without Losing the Thread

The old pattern was: ask the model to summarize a paper, then separately ask it to write code based on that summary, then separately debug the code, then separately interpret the results.

Each handoff was a place where context got lost and I had to re-explain what mattered.

With Opus 5.5, I gave it the question and the tools once.

It read the papers, noted that the speculative-decoding paper's reported speedup assumed a draft model roughly 20x smaller than the target model — a ratio we don't actually use in production — and it flagged that as a reason our real-world numbers might not match the paper before it even ran anything.

That's the kind of caveat a first-year researcher misses and a good senior one catches. It caught it unprompted.

It Writes the Experiment, Not Just the Summary

For the KV-cache eviction comparison, it didn't just describe the two algorithms.

It wrote a benchmark harness that instrumented memory pressure at three different load levels, ran both eviction strategies against our actual model serving stack, and produced a table with latency and memory numbers side by side.

Then it wrote up the result in about four paragraphs that I could have dropped into a design doc with minor edits.

The whole thing — reading, coding, running, writing — took 47 minutes.

The equivalent task took me the better part of two days when I did it manually last year, and a chunk of that was just re-familiarizing myself with the eviction algorithms' actual implementations, not the interesting part of the work.

It Knows When a Paper's Claim Doesn't Transfer

The most useful moment was the request-coalescing task.

The paper we were evaluating showed strong gains, but Opus 5.5 pointed out — correctly — that the paper's workload was almost entirely uniform request sizes, while our actual traffic has a long tail of oversized requests that would blow past the coalescing window.

It ran a synthetic benchmark mimicking our real traffic shape instead of the paper's benchmark and showed the gains mostly evaporated.

That's not pattern-matching on the text. That's noticing a mismatch between an experimental setup and a deployment context, which is most of what separates a useful research read from a useless one.

The Reality Check: Where the Loop Still Breaks

Here's where I stop the hype train, because I watched it happen twice.

Article illustration

On the eviction comparison, Opus 5.5 quietly conflated "lower average latency" with "better for our use case" — it didn't weight tail latency the way we actually care about it, because I hadn't told it to, and it didn't think to ask.

A junior researcher who'd sat in three of our incident retros would have known tail latency is what pages people at 2 a.m. The model didn't have that context and didn't know it was missing it.

The second failure was subtler and more concerning: on the speculative-decoding task, it generated a plausible-looking chart with numbers that were close to right but not identical to what I got when I reran the same benchmark myself.

Not fabricated exactly — the code was real and it did run — but a sampling seed wasn't fixed, and the write-up presented a single run as if it were a stable result. It didn't flag the variance.

I only caught it because I ran the benchmark three times out of habit.

That's the actual risk with this generation of tooling. It's not that it hallucinates nonsense — it mostly doesn't anymore.

It's that it produces work polished enough that you stop checking it, right at the moment checking it matters most.

A sloppy human draft gets scrutinized. A clean, confident AI draft gets rubber-stamped. That gap is where bad decisions live.

It's also still bad at the actual generative step — proposing a genuinely novel hypothesis nobody's tried.

Everything it did well in my test was recombination and evaluation of existing ideas, done fast and done carefully. None of it was a new idea.

That's fine; most research work isn't new-idea generation either.

But if your job is specifically "come up with the thing nobody's tried," Opus 5.5 is a fast research assistant for that job, not a replacement for the person doing it.

What I'm Actually Doing With This Now

I'm not handing off research and walking away. I'm restructuring how I spend my hours on it.

The literature review and first-pass experiment design now go entirely to Opus 5.5. That's the part of research work that's mechanical and time-consuming but low-judgment — reading twenty papers to find the three that matter, writing the first version of a benchmark harness.

I check its source list against my own quick skim rather than reading everything myself first.

I fix the seed and rerun everything it reports before I trust a number. Given what happened with the speculative-decoding chart, I now explicitly prompt it to run multiple trials and report variance, and I still spot-check by rerunning myself.

This isn't optional if the result is going into a decision that costs engineering time.

I own the "does this matter for us" judgment call entirely. The model can tell you what a paper found. It can even tell you where a paper's assumptions don't hold.

It can't tell you which metric your team actually gets paged on, because that context lives in your incident history, not in the literature.

That's the part of the job that isn't going away, and honestly, it's the part that was always more interesting anyway.

My backlog moved faster this month than it has in two years. The mechanical middle of research work — the part that used to eat entire days — got compressed into an hour.

What didn't compress is judging what the result means and deciding what to do about it.

If a tool can do 80% of the reading-and-running loop I used to bill entire days to, what's left of the job actually looks more like a job I'd want.

Has anyone else handed real backlog work to Opus 5.5 yet, and did it catch the same blind spots I did?


Story Sources

YouTubeyoutube.com