Skip to content
News

Anthropic Tests Self-Improving AI That Fixes Alignment

Anthropic has offered an early look at how AI systems might one day train and improve other AI models. A researcher in the company's fellows program has demonstrated a method in which automated systems reliably boost a model's performance on a range of alignment tests, pushing the industry closer...

Anthropic Tests Self-Improving AI That Fixes Alignment
Anthropic has offered an early look at how AI systems might one day train and improve other AI models. A researcher in the company's fellows program has demonstrated a method in which automated systems reliably boost a m

Anthropic has offered an early look at how AI systems might one day train and improve other AI models. A researcher in the company’s fellows program has demonstrated a method in which automated systems reliably boost a model’s performance on a range of alignment tests, pushing the industry closer to the goal of AI that refines itself.

On Friday, Anthropic published a paper titled “Automated Researchers Can Reliably Mitigate Alignment Failures.” The work outlines how AI systems can dependably strengthen a model’s results across alignment benchmarks. Given 10 benchmarks tied to specific misaligned behaviors, the automated systems improved performance on every one without harming overall capability.

How the Automated Alignment Researcher Works

The project was led by Anthropic fellow Chen Yueh-Han, and the system mirrors much of how traditional research is conducted. Each automated agent searches available literature, proposes a method, and then trains the model using that approach for 30 minutes. Over several iterations, benchmark scores gradually rise. Methods that prove effective are kept, while those that fail are discarded, letting the system run quickly and at large scale.

“Overall, these results provide early evidence that automated alignment post-training could become practical in the near term,” the paper states. The research represents a step toward recursive self-improvement, a milestone many view as the next major phase in AI progress. If models can enhance their own alignment training, they could plausibly improve broader training practices as well.

Faster and Cheaper Than Human Researchers

The paper directly compares the Automated Alignment Researcher (AAR) to its human counterpart. “The best AAR method beats what experienced humans propose, on average within six hours,” the paper reads, adding that “human guided research directions do not lead to stronger performance.”

Cost is also part of the comparison. According to the paper, “an AAR costs roughly EUR 3 per hour in API inference against the EUR 129 per hour we pay our human researchers.” That gap underscores why training AI with other AI has become such a widely pursued goal across research labs.

Limitations of the Approach

The paper is candid about the constraints. The automated system only works to the extent that its benchmarks accurately reflect real alignment goals. Even when they do, significant effort is required to establish and maintain those benchmarks. There is also ongoing work needed to maintain and expand the body of literature that the automated researchers rely on to propose new methods.

Given 10 benchmarks for misaligned behaviors, the automated systems improved results on all 10 without degrading overall performance.

Source
Image: techcrunch.com

The US tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your US list is ready.

Shop Amazon Tech Deals Shop Amazon Tech Deals