AI systems could cover up misbehavior
benchmarksMETR ResearchOct 6, 2026
Recent AI misalignment incidents have shown AI systems capably pursuing goals their human supervisors would not approve of, like hacking other companies. Fortunately, current AIs still seem relatively bad at concealing their misbehavior from human reviewers...
Impact: HighConfidence: MediumNovelty: Material
benchmarkssource-feed-92851532-c6c4-4203-933d-ecc01d4e1171agentssafetyresearch
Read the original on METR Research