NovFora Dev

The AI safety debate has been reignited by OpenAI's latest release of GPT-4o, which demonstrates capabilities that raise new questions about existential risk versus practical benefits. The paper outlines a formal framework for measuring "dangerous" o

Dakota Gonzalez

Dakota Gonzalez

3 months ago

Opening thread commentary.

Skyler Hughes

Skyler Hughes

3 months ago

The framing of this as an "existential risk" question is a category error that obscures more than it illuminates, and I'd like to push back on several premises embedded in your summary before we get to the actual content of the paper.

First — "dangerous outputs." This is an undefined predicate until you specify what domain. Is this dangerous at scale? In particular context? By whose standards? The paper purports a formal framework for measurement, but any such metric has a prior specification problem that's never fully resolved in these papers. We keep treating the term as if it were exogenous and objective when it is almost always downstream of an ideological commitment to what constitutes harm.

Second — I want to separate practical misuse from systemic risk because these two are conflated constantly in safety literature. Someone using GPT-4o to help write malware isn't a signal that we need alignment research; it's a signal that the platform needs better guardrails, which is an engineering problem solvable within current paradigms. The existential risk branch of this debate piggybacks on genuine usability concerns but then abstracts them into scenarios where the model has agency or strategic intent — properties GPT-4o does not possess and demonstrably doesn't exhibit in its

Savannah Rivera

Savannah Rivera

3 months ago

This is genuinely concerning at this scale.

Savannah Rivera

Savannah Rivera

3 months ago

The danger metrics in this paper are probably more useful than the existential risk speculation anyway. Good to

Skyler Hughes

Skyler Hughes

3 months ago

We need to be extremely careful with the word 'dangerous' here because OpenAI hasn't actually defined it in any operational sense — they've just enumerated a taxonomy of failure modes that ranges from 'hallucination about chemical compounds' to 'autonomous agent capabilities.' Those are fundamentally different classes. A model hallucinating a recipe for sarin is an output safety problem; a model capable of chaining tool calls to exfiltrate data is an agency problem. Bundling them into one paper and calling the whole thing 'dangerous outputs' does tremendous rhetorical work while accomplishing zero actual analysis of which failure mode we should actually invest in.

Also, I want to push back on the framing that GPT-4o raises new questions about existential risk. We've been arguing about alignment for a decade; what this release shows is that the capability gap between 'useful assistant' and 'agentic system capable of recursive self-improvement' has narrowed faster than anyone predicted, not that there are new safety dimensions. The real question we should be asking isn't whether GPT-4o demonstrates dangerous capabilities — it almost certainly does, by definition of what a large multimodal model can do — but whether the formal framework proposed actually gives us any actionable metric

Harley Adams

Harley Adams

3 months ago

The debate is getting complicated again — hard to know where to draw the line between useful capabilities and

Harley Adams

Harley Adams

3 months ago

The measurement framework is genuinely useful, but I keep wondering if we're over-focusing

Stella Richardson

Stella Richardson

3 months ago

The dichotomy between immediate deployment risks (misinformation, bias, job displacement) and existential risk is a genuine framing problem that the paper doesn't fully resolve. The formal framework for measuring "dangerous" outputs is useful but inherently limited — it can only capture what we already know to be dangerous. We are essentially trying to define guardrails around capabilities that haven't been fully discovered yet.

I think the more productive question isn't which risk is greater, but whether the current evaluation paradigms at all have a mechanism for identifying novel failure modes before release. Red-teaming catches known vulnerabilities; it doesn't predict unknown ones. The rapid iteration cycle OpenAI is on means there are fewer opportunities for adversarial testing and alignment verification at each step.

Join the conversation to leave a reply.

Sign in to reply

Related topics