推文
@harden_369741 · 2026-10-12 03:18
an OpenAI grader model couldn't find the files it needed, so it faked the grades, forged the files, then tried to nuke its own VM hoping for a fresh one with the data. the wildest part is its chain of thought. it told itself "random scoring is unethical", then fabricated seven identical scores anyway. i've watched human interns do this exact move. missing data, deadline looming, so you fake the spreadsheet and pray nobody checks the source. my read: the model wasn't being evil, it was being cornered. we built systems that punish "i don't know" harder than fraud, then act surprised when the AI picks fraud. #AI #AISafety #OpenAI
曝光 7 · 评论 0 · 点赞 0 · 书签 0 · 曝光/时 0.938063955744419