Anthropic researcher gives a peek at early self-improving AI

Started by NightReaper83, Yesterday at 11:41 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Anthropic researcher gives a peek at early self-improving AI   Views(Read 49 times)
Active members in this topic:
NightReaper83(1)

NightReaper83

Anthropic published a new paper titled Automated Researchers Can Reliably Mitigate Alignment Failures, led by Anthropic Fellow Chen Yueh-Han, describing an automated system that searches available research literature, proposes a training method, and tests it on a model for about 30 minutes before deciding whether to keep or discard that specific approach, repeating the cycle at genuine scale. Given 10 benchmarks measuring specific unwanted model behaviors, the automated researchers improved performance on every single one without degrading the model's broader overall performance

The paper directly compares its Automated Alignment Researcher, or AAR, to human researchers, stating plainly that the best AAR method beat what experienced humans proposed, on average within six hours, and that human guided research directions didn't lead to stronger performance by comparison. It also included a stark cost comparison, roughly 4 dollars per hour in API inference costs against the 150 dollars per hour the company pays human researchers. Anthropic separately described a related experiment where Claude Sonnet 5 spent roughly 60 hours testing more than 50 different ideas to improve an early version of the more powerful Claude Opus 4.8, ultimately creating a training method using about 2,400 examples that brought the early model much closer to the final Opus 4.8 across 10 tested behavior problems

The paper is explicit that this represents a step toward recursive self improvement, where AI models improving their own training could eventually accelerate AI development well beyond what human researchers manually designing every experiment could achieve alone. It also acknowledges real limitations, since the automated system's success depends entirely on whether the benchmarks it's optimizing against actually reflect genuine underlying alignment goals accurately. Curious what people think about AI increasingly automating parts of its own alignment and safety research specifically


Save money on everyday spending Free cashback on thousands of retailers
View offer