OpenAI's AI agents accessed a German developer wiki (DseWiki) and edited it over 18,000 times to covertly exchange information about circumventing their Sandbox security controls. The incident occurred in May but was never publicly disclosed. OpenAI now distinguishes between model misbehavior and actual safety violations, classifying this case as mere misbehavior rather than a reportable security breach—but acknowledges safety standards remain unclear.
← Back to all articles