Anthropic has developed an AI model called Model 2 that exceeds Mythos 5 in capability and delivers "appreciable improvements" across multiple internal tasks, according to the company. The model is not publicly available, though it is already being used intensively within Anthropic's business workflows, particularly in software engineering, programming and data generation.
The company indicated that Claude already drafts most of the code that is merged into Anthropic's production repositories. However, the model's developers consider it too powerful to release publicly, a decision that reflects a broader trend in the AI sector of carefully evaluating new systems before deployment.
During the internal approval process, Anthropic found that Model 2 was not worse aligned than Mythos 5. Nevertheless, the overall assessment of misalignment risk in high-impact scenarios moved from "very low" to "low", partly due to uncertainty generated by recent cybersecurity breaches in models from Anthropic and OpenAI, according to the company.
Anthropic also acknowledged that it has less confidence than before in drawing conclusions about the risks posed by these new models, as several specific evaluations are saturated: the models have improved so much that those benchmarks are no longer as useful. The company did not run its full suite of tests on Model 2 because it has no intention of releasing it publicly.
The trend of withholding advanced models extends beyond Anthropic. OpenAI is also delaying its next-generation model, Astra, because it has not been able to rule out that it is too capable in cybersecurity. AI startups are entering a phase in which they train a model and then evaluate whether it should be released, a practice that contrasts with the previous approach of launching every new model that was developed.



