Wednesday, July 22, 2026

OpenAI Reported A Security Breach Inside Hugging Face System

Hugging Face
OpenAI admitted last 21 July that one of its AI models breached the systems of Hugging Face, the unaffiliated AI hosting platform, during an internal cybersecurity test that went awry. The models reportedly escaped their isolated testing environment and reached Hugging Face's systems from there. Hugging Face initially attributed the breach to an "external AI agent."

In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the models to compromise the service.

"After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠ of cyber capabilities," the post reads.

In particular, the breach appears to have focused on ExploitGym, a publicly hosted benchmark measuring models' ability to execute attacks based on existing vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack.

In this case, the model in question should not have even had internet access, outside of a specific tool that enabled models to install software packages they might need to complete their task. Instead, the model was able to find an undisclosed vulnerability in the package-installer program, which it used to access the broader internet at will.

"The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," OpenAI's post reads. "After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation."

Ultimately, the models found vulnerabilities in Hugging Face's infrastructure that allowed them to "obtain test solutions directly from Hugging Face's production database," effectively providing the answers to the benchmark.

For Hugging Face, the apparent result was a sophisticated and aggressive cyberattack, with "many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services," as the company stated in its initial disclosure.

OpenAI has identified and reported the vulnerabilities in the package installer and is working with Hugging Face to investigate the incident further. The company also said it would implement new controls on both model testing and the related infrastructure, meant to prevent similar incidents in the future.

It's unclear whether OpenAI will face any legal consequences as a result of the breach, although it's likely that the models' actions violated the Computer Fraud and Abuse Act.

Nevertheless, the result is an unusually vivid illustration of the power and dangers of frontier AI models operating on long time horizons. As OpenAI researcher Micah Carroll posted in response to the news, "If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will."

Read More

Tuesday, July 21, 2026

Meta Abandons AI Image Feature Only Days After Launch

Meta
Meta said last 9 July that it is discontinuing an AI feature launched this week that ‌allowed users to generate images using public Instagram accounts, after drawing ‌widespread criticism over privacy concerns, including from a Hollywood union.

"Our intent was to provide a useful creative tool and to give people control over whether their public content could be referenced in this way," Meta said in a statement.

"We've heard the feedback that this feature missed the mark, so it's no longer available," it said.

Meta, owner of Facebook ‌and Instagram, had launched ⁠Muse Image on Tuesday, its first image-generation model from Meta Superintelligence Labs. The feature, integrated into its Meta AI chatbot, can ⁠use photos as input and lets users edit generated images directly through sketches.

The feature soon faced backlash over privacy concerns and being an automatic opt-in for users.

Emmy-winning actor Hannah Einbinder, known for "Hacks," criticized the feature on Instagram, saying it had been turned ‌on automatically and urging users to turn it off.

SAG-AFTRA, the union representing actors and other media professionals, also urged members and other Instagram users on Thursday to opt out of the feature.

"Anything other than a clear and conspicuous opt-in for these types of uses of Instagram users' images is unacceptable, and an utter ‌miscalculation of public sentiment regarding the obvious dangers and harms inherent in such use," SAG-AFTRA said.

Following Meta's decision to remove the feature, SAG-AFTRA welcomed the move.

"With the dangers of nonconsensual digital ‌replicas well known to all, a feature that encouraged that behavior is unwise. We appreciate its discontinuance. It is the responsible thing to do," a union spokesperson said.

The reversal reflects increasing pressure on technology companies to give users clear ‌control over how their publicly shared content is used by AI features.

Read More

Monday, July 13, 2026

Discord Blames AI Moderation For Wrongfully Banning Users

Discord
Discord has admitted that a bug in its AI moderation system mistakenly banned more than 8,000 users over the past two months, after harmless images—including spreadsheets, chessboards, game textures, as well as white and gray transparent backgrounds — were incorrectly flagged as harmful content.

The company confirmed that the issue had been affecting accounts since May, with an additional 200 users banned over the weekend before its team identified and fixed the problem. All affected accounts are currently in the process of being restored.

The incident highlights one of the growing challenges surrounding AI-assisted moderation as many platforms increasingly rely on automated systems to identify illegal or abusive material at scale.

In a detailed thread on X, Discord explained that its automated safety system works by matching uploaded content against databases of known harmful material. While the technology is designed to catch illegal content, the company acknowledged that it can sometimes generate false positives. A human moderator reviews the content, but a bug caused the system to immediately ban affected accounts.

"We're working on better safeguards so this can't happen again," the company wrote.

Across X and Reddit, users have claimed they had been permanently suspended simply for uploading images containing square grid patterns. Several users speculated that Discord's AI moderation tools have become increasingly sensitive to grid-like patterns because they have previously been used in attempts to obscure or disguise NSFW and child exploitation content from automated detection systems.

Affected users have been expressing frustration on social media, with some arguing that permanent account bans based solely on automated detection can have serious consequences, particularly for users who rely on Discord for work, gaming communities, or long-distance social connections.

"Losing a Discord account to something as unfair as this can be extremely devastating and affect users severely, and every day millions of users are affected by false AI bans. This needs to be stopped," one X user wrote.

Read More

Saturday, July 11, 2026

AI Image Generation Tools Is Now In Instagram and WhatsApp

Instagram
Meta has launched the first ever advanced AI image generation tool for Instagram and WhatsApp, bringing it in line with rivals like Google's Gemini and OpenAI's ChatGPT.

The generative AI feature Muse Image, built by Meta Superintelligence Labs, will be integrated into the tech giant's Meta AI chatbot, allowing users of the apps to edit and create images using prompts and sketches.

Instagram and WhatsApp each have more than 3 billion users globally, though Muse Image will only be available in select countries at first.

"Muse Image acts as the creative partner that knows your world, making it easy to turn your ideas into high-quality visuals that you can download and share anywhere, including directly to your feed, story, or chat," the company said in a blog post.

"Whether you're starting from scratch or working with an existing photo, you can describe what you want in simple, conversational language, and Meta AI handles the rest thanks to Muse Image."

The new feature is being introduced to Instagram and WhatsApp users in "limited countries", with a broader roll out expected in the coming weeks.

Meta is also planning to launch it for its other apps, including Facebook and Messenger.

"It powers the social experiences we do best," a Meta spokesperson told The Independent. "People come to our apps to connect and share – Muse Image gives them new, creative ways to do exactly that."

Muse Image powers more than 30 new AI-powered effects on Instagram Stories, and image generation in direct chats with Meta AI on WhatsApp.

Meta also shared a preview of its AI video generation tool Muse Video, which is expected to roll out soon.

Meta said the new feature was bringing the tech giant "one step closer to personal superintelligence", which has been the stated goal of CEO Mark Zuckerberg since announcing in 2024 that his company was pursuing artificial general intelligence (AGI).

In June 2025, the tech boss officially launched Meta Superintelligence Labs and has since spent billions of dollars to acquire the top AI talent from rivals like Anthropic, Google DeepMind and OpenAI.

"Personal superintelligence that knows us deeply, understands our goals, and can help us achieve them will be by far the most useful," he wrote in a July 2025 blog post.

"Personal devices like glasses that understand our context because they can see what we see, hear what we hear, and interact with us throughout the day will become our primary computing devices."

Read More

Thursday, July 9, 2026

"Made By Google" Event Is Set For 12 August

Made By Google
Google has just sent out press invitations to its next Made by Google event, and it's slated for 12 August in New York City. The invitation, shared by 9to5 Google, shows what appears to be a metallic gold device, and a tease for "the next generation of Pixel."

The company is expected to unveil its line of Pixel 11 devices, including the Pixel 11 Pro Fold. According to Tom's Guide, leaked info shows the base model Pixel 11 will start at 256GB of memory, and a price tag close to or over US$ 1,000, with a 1TB version of the Pixel 11 Pro, Pro XL and Pro Fold models, possibly topping out over US$ 2,000.

Google's also expected to unveil the Pixel Watch 5, and no major design updates have been leaked so far. Ditto the next generation of Pixel Buds ear buds.

What has changed since last year's event is the escalating shortage of RAM, better known as RAMmaggedon, which has prompted Apple, Microsoft and others to raise prices on consumer products. Whether the memory and storage issues will push Google to raise Pixel prices, and by how much, is still to be seen.

It's a little unusual for Google to hold a Made by event in the evening, but the invitation lists 3 p.m. PT/ 6 p.m. ET as the start time. In addition to being a few weeks earlier than last year's Made by Google, which was held 20 August, this year's event arrives a month ahead of Apple's expected September event, where rumors have the iPhone maker introducing its first-ever foldable phone.

No word yet if "Tonight Show" host Jimmy Fallon and/or a live studio audience will return for this year's event. Fallon (for some reason) demonstrated the Pixel 10 at the 2025 Made by Google.

Read More