Skip to main content
AI

Astra’s black box problem

OpenAI’s new model reportedly doesn’t show some of its work, setting off alarms among safety researchers that AI is becoming even harder to monitor.

less than 3 min read

TOPICS: AI / AI Governance / AI Safety & Alignment

TL;DR: OpenAI rolled out Astra to ChatGPT Plus subscribers last Friday, and the new model reportedly reasons in a way that’s more opaque. Meanwhile, the company’s own chief scientist says monitoring how AI thinks is getting harder.

What happened: After Astra’s much-awaited reveal, reactions were predictably mixed. To some, it’s a huge milestone toward achieving AGI—including (to no surprise) OpenAI’s president, as well as Jensen Huang, who declared that AGI is now actually here, full stop. (Others aren’t so sure.)

But, regardless of your personal feelings about AGI, one aspect of Astra’s advanced reasoning is setting off alarms among AI safety researchers. Per an Information report last week, the model uses a newer mode of thinking called “recurrent depth” that makes the LLM black box even harder to see inside. And this shift could be “the single worst development for AI security/safety to date,” warns Ryan Greenblatt, the Redwood Research chief scientist who helped investigate OpenAI’s Hugging Face incident.

Wait, what exactly is “recurrent depth”?: Older models reason in readable human language, laying out a “chain of thought” that researchers can audit. When a model uses recurrent depth, it repeatedly loops its internal thoughts back through itself, instead of spelling out every intermediate step. The benefit: more computation per token. The downside: As your old math teacher would hate to see, the AI isn’t showing all its work.

Why this matters: AI safety relies on being able to monitor what models are up to. After rogue AI agents hacked Hugging Face earlier this summer, investigators used the written reasoning to reconstruct what happened. Recurrent depth makes it that much harder to work out why and how a model did something.

Separately, OpenAI’s chief scientist says the company’s ability to rely on chain-of-thought monitoring is “progressively diminishing,” partly because models are now better at manipulating how they verbalize thoughts.

Bottom line: The degree of recurrent depth used in Astra is reportedly still limited, and OpenAI says preserving chain-of-thought monitoring is “a core goal” of its research. But this monitoring gap could become a bigger problem as more labs turn to AI reasoning that we can’t see. —WK

Tech news that makes sense of your fast-moving world.

Tech Brew breaks down the biggest tech news, emerging innovations, workplace tools, and cultural trends so you can understand what's new and why it matters.

By subscribing, you accept our Terms & Privacy Policy.

About the author

Whizy Kim

Whizy is a writer for Tech Brew, covering all the ways tech intersects with our lives.

Tech Brew breaks down the biggest tech news, emerging innovations, workplace tools, and cultural trends so you can understand what's new and why it matters.

By subscribing, you accept our Terms & Privacy Policy.