AI researchers are quietly freaking out

So I sat down to catch up on this week’s AI news expecting the usual shiny toys, and instead I hit a wall of AI researchers basically saying “we might all be in trouble.” The shiny toys were there too. But the twist this week isn’t a feature drop. It’s the people building the smartest systems on earth admitting they don’t have a plan.

This all comes from a weekly roundup by Matt Wolfe, the creator behind Future Tools who watches AI news every single day so the rest of us don’t have to. He pulled together the releases and the drama into one video, and I want to break down the parts that stuck with me.

What’s new this week

Let’s start light, because a lot actually shipped:

  • ChatGPT Images 2.5 (OpenAI): The big upgrade is consistency. Matt showed how it keeps a person’s face, pose, and details steady across edits. There’s also a new “Sketch” feature where you doodle a rough drawing, hit check, and it turns your scribble into a realistic image. He even played a telephone game, letting the AI sketch him and then making that sketch realistic.
  • Meta Muse: A personal agent that takes actions for you. It connects to email, calendar, and apps, runs on a secure virtual machine, and keeps working after you close the app. Matt called it the easiest agent he’s ever onboarded. It climbed to the number two app in the US.
  • DeepSeek V4.1 Flash: Dirt cheap at around 27 cents per task, and it scores shockingly high on coding benchmarks. Though Matt noted his own tests didn’t quite match the hype.
  • Rapid fire: Apple’s iPhone 18 Pro and folding iPhone Duo, new Apple Watch audio features, Microsoft MAI-Image 2.6, ChatGPT learning to write in your voice, Suno V6 trained only on licensed music, Google’s Lyria 3.5 music model, and a DaVinci Resolve update that plugs Claude into your video editing.

The twist that stopped me cold

Here’s the unexpected part. A researcher named Jacob Coxin publicly resigned from Anthropic after three years of pre-training work at both OpenAI and Anthropic. His claim, which Matt highlighted: neither company is acting responsibly, and they’re “racing straight to self-improving super intelligence.”

Then it got heavier. Evan Hubinger, Anthropic’s alignment science lead, basically agreed. He said he personally thinks there’s a greater than 10% chance AI could kill all humans within the decade, and that Anthropic does not yet have a plan to solve alignment for super intelligence.

Read that twice. The company that believes it’s the only one responsible enough to handle this is also saying it has no plan and isn’t on track to get one.

And it’s not just Anthropic. Matt pointed to an essay called An Alien Mind from OpenAI’s chief scientist Jakub Pachocki, warning that recursive self-improvement (AI making the next AI smarter, over and over) could take off fast. His line that landed hardest: as these systems get more capable, the results get harder for humans to interpret.

Why I think this matters

I was genuinely unsettled watching this part. Not panicked, but unsettled. Here’s how Matt framed the two sides, and I think it’s fair:

  1. The skeptic’s view: It’s a scare campaign to juice an IPO. Matt even noted that physicist Sabine Hossenfelder said she was offered money to tell people AI will kill us. So that incentive is real and documented.
  2. The other view: These are respected researchers whose literal job is mapping worst-case scenarios. Living in that bubble, they genuinely believe it. Matt doesn’t think they’re paid shills.

There’s also a practical wrinkle the author raised from the OpenAI essay: the proposed defense against bad AI is good AI. We’d need powerful aligned systems to protect infrastructure from rogue agents in real time.

One more jaw-dropper. OpenAI says it solved a Navier-Stokes Millennium Prize problem, unsolved for roughly 90 years, using an internal model “significantly more capable” than the GPT-6 Astra we all have access to. So the models we’re impressed by are apparently the weak ones.

The real challenge here

The honest problem, as Matt laid it out, is that there’s no clean off switch:

  • Nobody trusts anyone enough to pause together. US labs, Chinese labs, everyone is locked in an arms race.
  • Dismissing the worried researchers as “doomers” shuts down people with valid concerns.
  • But spreading pure doom just scares folks and deepens the us-versus-them divide.

Matt’s take, and I agree with it, is that the middle path is boring but right: pour more money and brainpower into alignment research instead of only chasing shareholder value. The toothpaste is out of the tube. Solutions beat panic.

Pro tips from the roundup

  • 🧪 Struggled with ChatGPT image consistency before? Give 2.5 a second look, the edit stability is the headline fix.
  • 🤖 New to agents? Muse is reportedly the gentlest on-ramp. Just connect apps and let it suggest.
  • 💸 Try Muse’s subscription audit trick. Matt had it scan his accounts and list every AI tool he pays for. It found a painful pile.
  • 📝 Want ChatGPT to sound like you? The new writing-style feature learns your quirks from Gmail, Slack, and Drive.

Worth your time

This one’s a mix of cool and uncomfortable, and I think that’s exactly why it’s worth watching in full. 🎯 Check out Matt Wolfe’s complete breakdown for the nuance on the alignment debate and demos of everything that shipped.

Scroll to Top