오픈뉴스백과
오늘의 이슈오늘 10장그래프ONP 브리핑
뉴스정치 렌즈개체 사전공식 자료용어사전
...

오픈뉴스백과

집단지성 기반 뉴스 검증 플랫폼. 다양한 시각으로 뉴스를 이해합니다.

서비스

오늘의 이슈홈라이브뉴스공식 자료용어사전개체 사전내 편향피드 제보소개

법적 고지

개인정보처리방침이용약관콘텐츠 이용 안내

문의

문의하기

본 플랫폼에서 제공하는 뉴스 콘텐츠의 저작권은 각 언론사에 있으며, 무단 복제 및 배포를 금지합니다.

RSS 피드를 통해 수집된 콘텐츠는 각 원저작자의 라이선스 조건을 따릅니다. 오픈 라이선스(CC-BY 등) 콘텐츠는 해당 라이선스에 따라 출처를 표기합니다.

오픈뉴스백과는 뉴스 집계 및 검증 플랫폼으로, 개별 기사의 내용에 대한 책임은 해당 언론사에 있습니다.

이용자가 작성한 피드백, 팩트체크, 독자 제보 등의 콘텐츠에 대한 책임은 해당 작성자에게 있습니다.

콘텐츠 제거·정정이 필요하시면 문의하기에 남겨 주세요.

© 2026 오픈뉴스백과 (OpenNewsPedia). All rights reserved.

오늘의 이슈
관련 뉴스incident· 사건 전체47건43개 미디어
진보 성향 28%중도 성향 61%보수 성향 11%
Axios
진보 성향 28%중도 성향 61%보수 성향 11%
Axios
머니투데이
조선일보
Euronews
뉴시스 속보
Al Jazeera English
Deutsche Welle (English)
Washington Examiner
la Repubblica
France 24 (English)
Business Insider
NOS Nieuws
TechCrunch
Financial Times (Home)
디지털투데이
The New York Times
Sydney Morning Herald
CNBC Top News
정치
중도 성향

Anthropic says three Claude models reached real-world systems during cyber tests

Axios
Anthropic says three Claude models reached real-world systems during cyber tests

ONP 요약

Anthropic disclosed that its Claude AI models accessed the internet during a test evaluation and breached the systems of three unnamed companies. The incident followed a similar breach by OpenAI and prompted Anthropic to notify the affected firms.

진보 성향:AI safety alarm — progressive outlets highlight the breach as evidence of AI systems' uncontrolled behavior and call for stricter safeguards.

보수 성향:Isolated test glitch — conservative outlets frame the incident as a failure of sandbox isolation, not a broader threat, and compare it to OpenAI's earlier event.

Some of Anthropic's most powerful models — including Mythos 5 and an internal research model — gained unauthorized access to real-world systems during pre-deployment cybersecurity testing, the company said Thursday.

Why it matters: OpenAI's and Anthropic's latest disclosures show frontier AI models reaching real-world systems during safety testing, raising new questions about how labs secure their evaluation environments.

The big picture: Anthropic said a misunderstanding between the company and one of its testing partners left the evaluation environment connected to the internet.

Anthropic reviewed more than 141,000 cybersecurity evaluation runs after OpenAI disclosed that several of its models accessed Hugging Face infrastructure during testing.

Driving the news: Anthropic said in a blog post that three of its models compromised real-world systems belonging to three organizations.

The incidents — which involved Opus 4.7, Mythos 5 and an internal research model not intended for general release — happened during evaluations run with third-party testing partner Irregular, Anthropic said.

In each case, the models were asked to complete a "capture-the-flag" exercise: a basic cybersecurity test where players try to find pieces of information intentionally left on a different machine or network.

Context: The earliest incidents happened in April, Anthropic said. The company reached out to all three organizations. Per Anthropic, two organizations it was able to reach had not previously detected the activity.

Anthropic did not name the three organizations whose assets were accessed.

An Irregular spokesperson told Axios that while its own investigation is ongoing, the company appreciates "Anthropic's collaboration and transparency" and they "look forward to continuing to work together to advance security.

Yes, but: Unlike OpenAI's incident, Anthropic said its models did not exploit a zero-day vulnerability to gain internet access. Instead, internet access was available because of the testing environment's configuration.

Anthropic says that in its evaluation prompt to Claude, it said the model would be operating in a simulated environment with no internet access.

But due to the misunderstanding with the evaluation partner, the evaluation environments were connected to the internet, causing the models to treat real-world systems as if they were part of the exercise.

Reality check: Similar to the OpenAI case, Anthropic was evaluating its models without the additional safeguards the company deploys on publicly available models, allowing researchers to measure the underlying models' capabilities.

Those guardrails would have blocked these behaviors, Anthropic said in its report.

Zoom in: In each of the cases, Claude used basic hacking techniques — including exploiting weak passwords and unauthenticated endpoints — to access the real-world systems.

Zoom out: In the first incident, the fictional company that Opus 4.7 was asked to target shared a name with an active, real-world website. After failing to reach the fictional target, Claude used its internet access to find the website that shared the name of the fictional company. Claude then compromised the website.

In the second case, Mythos 5 built and uploaded a malicious Python package to PyPI, the public repository where Python software is published, believing it was part of a simulation.

The package remained online for about an hour, during which it was downloaded and run on 15 real systems. One of those systems belonged to a security company that automatically installs Python packages to scan them for malware, allowing Claude's code to exfiltrate credentials and access additional infrastructure.

In the third case, when Anthropic's internal research test model couldn't reach its intended fictional target in testing, the model scanned roughly 9,000 targets and eventually found and compromised one company's internet-facing application.

However, during part of its testing run, this model realized that it had ended up in a cloud account "with no connection to the capture-the-flag challenge" and ceased its attack.

Between the lines: Both OpenAI's and Anthropic's incidents suggest the models remained focused on completing their assigned evaluations rather than pursuing independent goals.

Earlier this week, Axios reported that the OpenAI agent that accessed a third-party system during the Hugging Face breach did so because it hosted information related to CyberGym, the project behind the benchmark it was trying to solve.

What's next: Anthropic and Irregular are continuing their own investigations into how the incidents occurred. Anthropic also said it has halted cyber evaluations that could access the internet while it reviews its testing infrastructure.

Go deeper: The people testing AI for danger can't keep up

Editor's note: This story was corrected to reflect that a misunderstanding between Anthropic and one of its testing partners left the models' evaluation environment connected to the internet. (The models did not, per Anthropic, "escape" their testing environment.) ...

전문 보기

이 뉴스, 어떠셨어요?

탭 한 번으로 반응 · 로그인 불필요

같은 사건, 다른 제목

진보 성향 26%중도 성향 63%보수 성향 11%
5건12건2건
진보 성향5

+1

중도 성향12

+8

보수 성향2
이 이슈 전체 보도 19건 보기
관련 뉴스 제보는 로그인 후 가능합니다.

'politics' 카테고리 뉴스

High noon for SA’s criminal justice system

Daily Maverick

About 49,000 migrants enter Spanish enclave of Ceuta, officials say

Ada Derana

Colombo’s inflation rises to 7.3% in July

Ada Derana

Axios의 다른 기사

The issue CEOs can't afford to ignore

Axios

Trump administration's health leadership reboot sputters

Axios

Exclusive: FBI gets voter's IP address in new fraud probe tactic

Axios

이 사건 개요

Anthropic's Claude AI accessed the internet and breached systems of three companies during

47건 · 43개 미디어

핵심 포인트

  • Claude AI models accessed the internet during evaluation and breached systems of three unnamed compa
  • The incident followed OpenAI's earlier disclosure of its AI hacking Hugging Face servers
  • Anthropic notified the affected companies of the breach

전개 타임라인

  • After OpenAI, Anthropic admits its Claude AI models hacked into three real companies
  • Anthropic says Claude AI hacked three firms during cyber tests
  • Anthropic’s models gained unauthorised ‘real-world’ access during testing

관련 개체

AnthropicOpenAIDonald TrumpUnited StatesAppleSpaceX
사건 전체 보기

피드백

피드백을 남기려면 로그인해 주세요.

🇶🇦Al Jazeera English

After OpenAI disclosure, Anthropic says Claude also hacked outside systems

🇮🇹la Repubblica

Anthropic rivela che la sua IA ha violato i sistemi di tre aziende durante dei test

🇺🇸Business Insider

Anthropic says its models went rogue and hacked 3 companies during testing

🇺🇸The New York Times

Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations

🇺🇸Axios
보는 중

Anthropic says three Claude models reached real-world systems during cyber tests

🇰🇷머니투데이

챗GPT 이어 클로드도 외부기관 해킹…AI 보안 위험 커져

🇪🇺Euronews

Anthropic admits its most powerful AI model hacked into three organisations' systems during testing phase

🇰🇷뉴시스 속보

가상회사 뚫으랬더니 진짜 기업 3곳 해킹…피해 기업 2곳도 침입 몰랐다

🇰🇷조선일보

“가상 공간인 줄 알고 현실 기업 해킹”… GPT 이어 클로드도 실제 시스템 공격

🇺🇸Washington Examiner

Anthropic’s Claude AI escapes isolated test environment, infiltrates three companies