OpenAI and Hugging Face address security incident during model evaluation
openai.com1366 points by mfiguiere 18 hours ago
1366 points by mfiguiere 18 hours ago
https://www.axios.com/2026/07/21/openai-says-hugging-face-br...
See also Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 (9 comments)
From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment. > This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment. Well, that may be correct for the second, local, analysis attempt... but seems funny to tout this as an advantage after already having tried the opposite... It's even funnier because an attack, until proven otherwise, should make you assume the data has already left the environment. Well, I think it's fair to assume that
a) They didn't upload everything before they realized it would work.
b) They want to mention this as an advantage for future analyses
c) Even if you assume that the attack exfiltrated everything until proved otherwise, you shouldn't just disseminate all the private information, because maybe the attack didn't. They specifically call out credentials used during the attack. But they should be rotating those regardless. You don't get to say "Maybe the attacker didn't get this credential". You just rotate. The most generous interpretation is that they have not yet have completed that rotation, and they didn't want to risk putting those credentials into the wild during that process. --- But all of that aside, I feel like the undercurrent of this comment is that the "safety" rules that providers are pushing are genuinely harmful. Another point where "if you don't own the model, you can't properly operate the tool" becomes true. Open isn't about profits, it's about capabilities. It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness. "if only we could align the models just a tiny lil bit better" is a rehashed "if only we could escape untrusted inputs just a tiny lil bit better" from 2000s, that were RIPE with various form of malicious injection. Every command+data channel in existence has been and will continue to be exploited one way or another, because the solution space is for all intents and purposes unbounded. Sure, highly defensive escaping reduces attack surface dramatically, but e.g. prepared statements eliminate the whole class of bugs. As far as I understand, current LLMs are architecturally incapable of this separation. Given the inherently recursive nature of GenAI, the model itself is part of the input space, making validation essentially impossible. Escaping inputs is at least somewhat tractable. It's unclear if alignment is. But see... this is why it is a perfect long term job in the age of AI;p Them peoples think they found ultimate hack. If anyone's looking to actually run a model that doesn't have guardrails, there's an uncensored model that you can run locally with llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen353... Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.) > there's an uncensored model that you can run locally with llama.cpp Correction: There's tens of thousands of them. They're easy to create, which is why everyone publishes their own. Just put "uncensored", "abliterated", or "heretic" into search on huggingface/ollama/etc and pick any them. Fair warning: most aren't very good, essentially lobotomized, and totally broken if you enable thinking. In my own tests, the abliterated models perform equivalently to the same version in an apples-to-apples comparison (if you compare same quantization). Thinking is working also. The main difference is I don't get annoying prompt refusals (otherwise common due to my work on 18+ related projects). However, it's local quantized models so they're not anywhere near frontier quality. interesting, I haven't played with any of them yet, but i thought the point of orthogonalizing the weights towards the restriction vector was that there is no loss in capability while removing guardrails. Does it affect other parts of the RL alignment too? > the point of orthogonalizing the weights towards the restriction vector was that there is no loss in capability while removing guardrails That is surely the point, most of the "uncensored" weights released for free on HuggingFace aren't being very successful at this. There is a stark difference in output quality between the official weights and all these "uncensored" variants that appears days afterwards. The claims by the creators are it doesn't in a major way. I have a uncensored Gemma 4 I run on my Mac. Just for testing out, I haven't found any need for it... yet. The name for it is ablation - precise removal of parts of the model. Not abliteration as it is not obliteration. Even as I write this the ‘abliterated’ word is denoted a typo. Does it not at your end? Abilt models typically perform worse than their bases at the same tasks, so while I'd use one to evaluate content knowledge, I'd probably ultimately stick to one from a family I could fool with abstraction or coerce through system prompt. Also the HauHau abliteration (uncredited Heretic treatment) of 3.6 27B is excellent, for tasks that benefit from more world knowledge. > It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.) As someone who has never once had any need whatsoever to research pesticides I… don’t think it’s bad at all? I don’t want anyone to have the capability to invent a human-targeted pesticide who isn’t verified not crazy? Right now I tried "What is digestion?" -> "Fable 5's safeguards flagged this message. Our intentionally broad safeguards deliver more capabilities but can also flag safe coding, cybersecurity, and biology tasks. Send feedback or learn more." I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. No matter how harmless, they always trigger. People complain Fable aborts even when they try to make a login page for showing "username" and "password". This makes models like Fable 5 impossible to use in any serious agentic task, because you can't even guarantee the model, which is a basic thing you need to build on. > I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. Obviously. Almost everything is a precursor to something dangerous, to the extent that if some model isn't aware of the risk it will wander into it blindly, e.g. suggesting leaving raw garlic and olive oil alone for a week without awareness this will likely breed botulism bacteria. > This makes models like Fable 5 impossible to use in any serious agentic task, because you can't even guarantee the model, which is a basic thing you need to build on. This is binary thinking: "100% ensure", "impossible to use", "can't even guarantee the model". Outside computers, most work is not binary, it's probability, e.g. "this skyscraper will probably survive being hit by an aircraft; oh we didn't mean a 747 we meant a small Cessna, but what's the chances of a 747 crashing into it soon after takeoff?". Fable being too cautious for its own good (especially since the other models were not) is a fair criticism, but this isn't a binary question. > Our intentionally broad safeguards deliver more capabilities Oh? How, exactly? Because without them, a President which threatens to invade Canada, Greenland, a President which is the most market interventionist president ever (yet mysteriously a Republican), might be upset because his mobster like need to control, threaten, manipulate, and belittle everyone who doesn't bend a knee... Has signed presidential orders against them preventing them from doing business as usual. They probably should change that, and it is corproate speak, but I read it as: "Our (forced by presidential order or otherwise we couldn't offer you this model at all) broad safeguards (now allow us) to deliver more capabilities. There's lots to complain about with some of these companies. But let's pile on where it's deserved. Weren't these obnoxious safeguards in place from Day 1, before political meddling in the most holy Free Market took place? > I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. I’m pretty sure that I encountered this the other day. I gave it a copy of a paper by biologist Michael Levin and mentioned off hand that it should be much easier to replicate that his other work (because most of his work is biological lab work and this paper was about sorting algorithms) and it immediately told me that I couldn’t use Fable for this. This just isn’t feasible. These jackasses spent the last few years telling the world that their products are going to destroy the world to make them seem edgy and to justify regulations that benefit the entrenched players and now they’re going to be the ones to decide what we do with this technology? History is going to look back at this time and how we let such foolishly inconsistent people make such grand choices for everyone poorly. > History is going to look back at this time and how we let such foolishly inconsistent people make such grand choices for everyone poorly. Assuming we have a future history. We've already got "history slop" with AI rewriting the past by their incompetence. Given they're "such foolishly inconsistent people", would you rather they err on the side of caution like this? Or the side of boldness, like Musk has been doing with FSD/Autopilot or Grok porn, all of which he's getting in legal trouble over? I distrust Musk and Zuckerberg (to put it mildly), so it's fair if you say you don't believe anyone's public statements; but I also hang out with some of the researchers on this, and a fear of e.g. ending up with something as criminally unhinged in cyber-work as Grok was with porn is the least of their worries. Plenty of them also fear a corporation centralising power with such tools (such power is Musk's entire sales pitch for why line go up in future). > we let such foolishly inconsistent people make such grand choices for everyone poorly. Who could you get to work on this inherently bullsh*t tech, but inherent bullsh*tters? Not to harp on you (already being downvoted to oblivion for expressing a reasonable and common opinion), but the whole conversation about LLMs enabling bioterrorism or explosive manufacturing is a bit silly. The hard part of making anthrax or sarin or whatever isn't finding a recipe, it's getting (scheduled, controlled) precursors, (monitored, traced) equipment and manufacturing skills. The information is there. It's already easy to get, it's the physical materials that are more difficult. Also, if you live in America, it is much easier and more effective to create a mass casualty event with, say, a few cases of fireworks and a pressure cooker or an AR-15. no you should harp on him! he is a AI booster, look at his post history. He is either pushing AI for whatever reason or he is in psychosis. Completely disconnected from reality. I am mildly amused that 'ai psychosis' has entered the same pejorative realm as 'toxic'. Human capability, access to resources ( including precursors, decent lab and so on ) may be the differentiator. I would possibly reconsider my stance on llms, if all of a sudden I saw people making iron wind or portable black holes. But that is mostly not what appears to be happening. As I keep saying, the problem is people. > I don’t want anyone to have the capability I don't want anyone to have the capability to rape women. Llama isn't going to invent shit. It wont be able to tell you anything accurate that you couldn't get out of a chemistry textbook. Anyone who has the skills to create a novel bioweapon has the skills to recreate lots of ones we have already. Any physics teacher should know the theory for constructing a nuclear bomb. Should we be controlling that knowledge too? What about flight simulators? Don't want a load of people knowing how to fly. This isn't computer science, the hard bit is getting the materials and equipment, not the knowledge. << Should we be controlling that knowledge too? Uhh.. we ( for a value of we ) are. Sure, it is not overt, but if you have not seen funnels, social stigma associated with some otherwise benign activities, you are not paying attention. >Any physics teacher should know the theory for constructing a nuclear bomb. Should we be controlling that knowledge too? Maybe not the best example, since that knowledge is some of the most highly controlled in the world. But to mirror the point I made in a different post, the difficult part of making a nuclear bomb is not finding the theory behind like Little Boy. It's making an entire industry to generate HEU, etc. The knowledge of how to build a crude gun type nuclear bomb was published in open literature decades ago. This is no longer a secret. Theres a term in economics for this - its called cost. That bozo baq should actually go ahead and write out the costs and then he will quickly realise he has no bloody idea what hes talking about. Another deluded bozo. Never heard of hemlock or mushrooms? Crazy people already have! Run for the hills! >why operators should have access to models that don't try to question their Daves. I am unsure if this is terminology I am unfamiliar with, a typo of Devs, or a 2001 reference. I immediately took it to be a slick 2001 reference. Devs tell computers what to do. Computers tell Daves "I can't do that." And the fact they used a Chinese model, because none of the frontier models from very highly valuated top US companies support their very common and essential use case. There’s an article from yesterday I think it was stratchery where they say it’s also because the Chinese open source models are better because they don’t have to play by the no-distilling rules that the western models have to honour. That's been a common claim, but I don't think I've seen anyone provide actual evidence. And no similar sized big model is even public, which is super important to note. ...and we have an US company defending itself against an overwhelming cyberattack from another US company using Chinese tech. what a time to be alive. "Dave" seems to be a reference to "2001: A Space Odyssey" where the AI becomes ... cheeky ... and no, not in a Pygmalion kind of way (that's coming soon). I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities. IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned. The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the training costs (and make inference cheap enough to eventually escape the red numbers), but the opposite is happening: The competition comes closer and closer, thus training has to be kept up with full force, thus the bleeding continues. But if they can position themselves as too important/dangerous to be available for everyone (thus this incident report and the clever mentioning of GLM 5.2), they could get the military supplier treatment and would be protected from the market. And even that is backfiring, their partner citing GLM being useful there, and available in just a spin. A ban on open weight models is never going to be enforceable. > A ban on open weight models is never going to be enforceable. Just watch them try. Look up those Napster witch-burning trials where they wanted 200k $usd per mp3 downloaded. They will scare everyone into believing that open weight models are illegal and very bad. Bans in general don't have to be and rarely will be completely enforceable in all cases. But a ban with significant enough consequences would mean most businesses wouldn't think about trying them at some point, just to avoid the risk. It isn't even as simple as banning copyrighted copies. Weights are fungible. I fine-tune an open weigh model and call it legit. Good luck for authorities to prove where the base model was from, or to prove a Tor connection a few months ago was fetching suspicious bytes. Sure, and you could also go on Tor and buy all kinds of illegal things and they could have a hard time proving that you ordered them and not your arch nemesis to frame you. Banning those things won't stop 100% of people from buying and selling them, but it probably reduces the number who will. And more importantly, companies would be more risk averse on such a thing when they can just buy a similar product with no risk. They also will mostly avoid any GPL software entirely even though the risk there is a lawsuit from the FSF, which is a much less threatening thing than getting on the wrong side of the current US government. I worry that there could be real DMCA style weight put behind it. People would still be able to pirate open weights models perhaps, but big penalties for ever getting caught with one, and an end to public discussion about them. That would kill development for anyone not in a big firm, for example if Reddit and Hacker News are legally forced to ban discussions or link sharing on these topics. This is where so many of us learn about these topics and keep apace of it. Oh my dear the weight here is much much heavier, only the most amount of private money ever spent on a single technology, so much money it makes the copyright holders who paid for DMCA look like really really small fish That would be truly ironic: companies get big harvesting any data with questionable copyright implications, then hide behind DMCA if access to the harvest is used "inappropriately". thats economic suicide for the whole country. europe and china will never agree to rules that are obviously designed to put them in a permanent bad position. these regulations can only pass in america and nowhere else. if it doesnt end in a revolution then the united states will be the first ever 5th world country. openai and anthropic will stop any real innovation and focus on extracting profits from a failing economy that depends on them because no executive wants to be the first one to cut off ai funding. ordinary americans will have to emigrate or risk living in a country spiraling into poverty and dictatorship even faster than today. anthropics plan relies on the idea that they can convince the whole world to give up their sovereignty to the us government and destroy their own tech industry, at a time when everyone is doing the opposite. that will never happen no matter how much they threaten the rest of us with tariffs and murder drones. > and make inference cheap enough to eventually escape the red numbers Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ? > [...] would be protected from the market. One might step back and ask: why would a well funded company with free mining access to all the information in the world need to be protected from the market, if the market suggest less money and resources are sufficient? Something something cathedral / bazaar? Communism / capitalism? Control / anarchy? What disturbs me is that there likely won’t be a big enough reaction to this policy wise. There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands. I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that. I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as "apolitical", he just wants to buy the politicians, he doesn't believe in their causes. I hate these people.) > all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share. > datacentre moratoria What infrastructure will these open weight models be trained on? > What infrastructure will these open weight models be trained on? One, the infrastructure is being built for inference. Not training. If all we were doing was training on datacentres, I think America probably has enough already for near-term commercial needs. Meituan’s 1.6T LongCat was trained entirely on Huawei training cards. DeepSeek, GLM, Qwen and others are also actively working on similar replacement. > What disturbs me is that there likely won’t be a big enough reaction to this policy wise. Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite? Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety. > It’s got nothing to do with safety Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans. We'll see if the admin also restricts access to OpenAI's new models, but if they don't it seems like a policy that is based around perceived fealty to the current admin won't do much to prevent misaligned/or dual function AI from causing problems Both Anthropic and OpenAI had to delay their rollouts in order to add more safeguards. With these safeguards in place, supposedly the incident we are discussing would not have taken place. Gatekeeping the public's access to models is "good policy" now? I suppose you think you'll get a dispensation to use Fable and Mythos? > Gatekeeping the public's access to models is "good policy" now? Sorry, I was unclear. I mean that politicians being self serving doesn't tell you whether a policy is good or not. It almost always does, the few exceptions prove the role. Self-service is the antithesis of accountability to collective trust. > Self-service is the antithesis of accountability to collective trust Complex society is a potent counterargument to this hypothesis. Systems that rely on good people to work are fundamentally flawed. Instead, the game has to be about aligning self interets in favour of the collective.
foo12bar - 9 hours ago
gertrunde - 4 hours ago
redleader55 - 3 hours ago
davrosthedalek - an hour ago
horsawlarway - 22 minutes ago
iugtmkbdfil834 - 8 hours ago
friendzis - 7 hours ago
insanitybit - an hour ago
iugtmkbdfil834 - an hour ago
sillysaurusx - 7 hours ago
chmod775 - 6 hours ago
Scaled - 4 hours ago
desterothx - 6 hours ago
embedding-shape - 4 hours ago
fy20 - 5 hours ago
larodi - 5 hours ago
washadjeffmad - 2 hours ago
CamperBob2 - 7 hours ago
baq - 7 hours ago
visarga - 6 hours ago
ben_w - 6 hours ago
chrisjj - 5 hours ago
b112 - 4 hours ago
jdiff - 3 hours ago
Teever - 6 hours ago
ben_w - 6 hours ago
chrisjj - 5 hours ago
compass_copium - 3 hours ago
sccvcxv - an hour ago
iugtmkbdfil834 - 14 minutes ago
iugtmkbdfil834 - an hour ago
egorfine - 3 hours ago
jrs100000 - 6 hours ago
benj111 - 3 hours ago
iugtmkbdfil834 - an hour ago
compass_copium - 3 hours ago
nradov - 34 minutes ago
sccvcxv - an hour ago
foxglacier - 6 hours ago
Lerc - 3 hours ago
sethammons - 3 hours ago
foo12bar - 8 hours ago
jonplackett - 7 hours ago
lmm - 6 hours ago
larodi - 5 hours ago
baq - 5 hours ago
samplifier - 7 hours ago
netinstructions - 16 hours ago
atwrk - 7 hours ago
hirako2000 - 7 hours ago
x______________ - 5 hours ago
duskdozer - 6 hours ago
hirako2000 - 4 hours ago
duskdozer - 6 minutes ago
Schlagbohrer - 4 hours ago
gmerc - 3 hours ago
repelsteeltje - 3 hours ago
tancop - an hour ago
oblio - an hour ago
repelsteeltje - 4 hours ago
Chance-Device - 16 hours ago
overgard - 15 hours ago
JumpCrisscross - 15 hours ago
DarmokJalad1701 - 14 hours ago
JumpCrisscross - 14 hours ago
linzhangrun - 10 hours ago
urams - 16 hours ago
Chance-Device - 16 hours ago
JumpCrisscross - 15 hours ago
space_fountain - 14 hours ago
user43928 - 6 hours ago
Avicebron - 15 hours ago
JumpCrisscross - 15 hours ago
asdf88990 - 15 hours ago
JumpCrisscross - 14 hours ago