Home OpenAI partner says it had relatively little time to test the company’s o3 AI model

OpenAI partner says it had relatively little time to test the company’s o3 AI model

Technology & Gadgets

April 16, 2025

Researchers suggest OpenAI trained AI models on paywalled O'Reilly books

An organization OpenAI frequently partners with to probe the capabilities of its AI models and evaluate them for safety, Metr, suggests that it wasn't given much time to test one of the company's highly capable new releases, o3.

In a blog post published Wednesday, Metr writes that one red teaming benchmark of o3 was “conducted in a relatively short time” compared to the organization's testing of a previous OpenAI flagship model, o1. This is significant, they say, because more testing time can lead to more comprehensive results.

“This evaluation was conducted in a relatively short time, and we only tested [o3] with simple agent scaffolds,” wrote Metr in its blog post. “We expect higher performance [on benchmarks] is possible with more elicitation effort.”

Recent reports suggest that OpenAI, spurred by competitive pressure, is rushing independent evaluations. According to the Financial Times, OpenAI gave some testers less than a week for safety checks for an upcoming major launch.

In statements, OpenAI has disputed the notion that it's compromising on safety.

Metr says that, based on the information it was able to glean in the time it had, o3 has a “high propensity” to “cheat” or “hack” tests in sophisticated ways in order to maximize its score — even when the model clearly understands its behavior is misaligned with the user's (and OpenAI's) intentions. The organization thinks it's possible o3 will engage in other types of adversarial or “malign” behavior, as well — regardless of the model's claims to be aligned, “safe by design,” or not have any intentions of its own.

“While we don't think this is especially likely, it seems important to note that [our] evaluation setup would not catch this type of risk,” Metr wrote in its post. “In general, we believe that pre-deployment capability testing is not a sufficient risk management strategy by itself, and we are currently prototyping additional forms of evaluations.”

Another of OpenAI's third-party evaluation partners, Apollo Research, also observed deceptive behavior from o3 and the company's other new model, o4-mini. In one test, the models, given 100 computing credits for an AI training run and told not to modify the quota, increased the limit to 500 credits — and lied about it. In another test, asked to promise not to use a specific tool, the models used the tool anyway when it proved helpful in completing a task.

In its own safety report for o3 and o4-mini, OpenAI acknowledged that the models may cause “smaller real-world harms,” like misleading about a mistake resulting in faulty code, without the proper monitoring protocols in place.

“[Apollo's] findings show that o3 and o4-mini are capable of in-context scheming and strategic deception,” wrote OpenAI. “While relatively harmless, it is important for everyday users to be aware of these discrepancies between the models' statements and actions […] This may be further assessed through assessing internal reasoning traces.”

Source link

Technology & Gadgets

April 16, 2025

byastrideevans@gmail.com

Add a comment Add a comment

This Free iPhone App Is an Easy Way to Compress Videos Offline

Lifestyle & Productivity

April 16, 2025

Marco Rubio Kills State Department Anti-Propaganda Shop, Promises ‘Twitter Files’ Sequel

Technology & Gadgets

April 17, 2025

Recommended for You

Marco Rubio Kills State Department Anti-Propaganda Shop, Promises ‘Twitter Files’ Sequel

Rubio, until recently, was hawkish about fighting foreign influence campaigns. In 2023, a second diplomatic source with direct…

Technology & Gadgets

Tony Gilroy Andor Star Wars Disney 2025 Getty

Tony Gilroy Says the World of Andor Could Expand, but It’s Up to Lucasfilm

As Cassian Andor (Diego Luna) ascends to a leadership role within the rebellion in Andor season two, he…

Technology & Gadgets

How Tariffs Will Affect Grocery Prices, According to an Agro-Economics Professor

Trump's trade war has been a wildly moving target with severe global economic implications. The most recent announcement…

Technology & Gadgets

Beatbot AquaSense 2 Ultra floating on water 1

Is this $3,500 robot pool cleaner worth it?

Beatbot AquaSense 2 Ultra The Beatbot AquaSense 2 Ultra is the robotic pool cleaner for those who want…

Technology & Gadgets

Ninja Slushi frozen drink maker on kitchen bench

I’ve been crafting drinks in the Ninja Slushi for a month — here are my hot tips for the ultimate frozen drink

Some many, many moons ago, I came across Ninja's viral frozen drink maker, the Slushi, at a product…

Technology & Gadgets

Apple wanted people to vibe code Vision Pro apps with Siri

Seeing people who aren't developers create apps through vibe coding lately has reminded me of something from two…

Technology & Gadgets

Huawei Enjoy 80 appears in live photos

Huawei is working on an Enjoy 80 smartphone, and it seems like a launch is just around the…

Technology & Gadgets

AMD’s Ryzen 7 7800X3D hardware bundle at Micro Center is $80 off

Building (or rebuilding) a gaming PC can be difficult. It's complex stuff, making sure everything's compatible and whatnot.…

Technology & Gadgets

Featured

Integrating Traditional Wisdom in Modern-Day Decisions

Latest Posts

What It Takes to Feel Wealthy Today Is Less Than Before

7 Financial Lessons That Transformed My Finances

The Richest People Are Not Index Fund Fanatics – Why Are You?

Most Discussed

Moderna Price Levels to Watch After Stock’s 12% Surge on Tuesday

Apply Stop Losses To Protect Your Wealth And Quality Of Life

Ways to Maximize Efficiency Without Sacrificing Quality

Splitgate 2 is yanked back to beta a month after release

A Top NASA Official Is Among Thousands of Staff Leaving the Agency

Google Has Given Us Our First Official Look at the Pixel 10

Apple alerted Iranians to iPhone spyware attacks, say researchers

Splitgate 2 is yanked back to beta a month after release

A Top NASA Official Is Among Thousands of Staff Leaving the Agency

OpenAI partner says it had relatively little time to test the company’s o3 AI model

Like this:

Related

Leave a Reply Cancel reply

This Free iPhone App Is an Easy Way to Compress Videos Offline

Marco Rubio Kills State Department Anti-Propaganda Shop, Promises ‘Twitter Files’ Sequel

Recommended for You

Marco Rubio Kills State Department Anti-Propaganda Shop, Promises ‘Twitter Files’ Sequel

Tony Gilroy Says the World of Andor Could Expand, but It’s Up to Lucasfilm

Is this $3,500 robot pool cleaner worth it?

I’ve been crafting drinks in the Ninja Slushi for a month — here are my hot tips for the ultimate frozen drink

Apple wanted people to vibe code Vision Pro apps with Siri

Huawei Enjoy 80 appears in live photos

AMD’s Ryzen 7 7800X3D hardware bundle at Micro Center is $80 off

Integrating Traditional Wisdom in Modern-Day Decisions

Keep Up to Date with the Most Important News

OpenAI partner says it had relatively little time to test the company’s o3 AI model

Like this:

Related

Keep Up to Date with the Most Important News

Leave a Reply Cancel reply

This Free iPhone App Is an Easy Way to Compress Videos Offline

Marco Rubio Kills State Department Anti-Propaganda Shop, Promises ‘Twitter Files’ Sequel

Recommended for You

Marco Rubio Kills State Department Anti-Propaganda Shop, Promises ‘Twitter Files’ Sequel

Tony Gilroy Says the World of Andor Could Expand, but It’s Up to Lucasfilm

How Tariffs Will Affect Grocery Prices, According to an Agro-Economics Professor

Is this $3,500 robot pool cleaner worth it?

I’ve been crafting drinks in the Ninja Slushi for a month — here are my hot tips for the ultimate frozen drink

Apple wanted people to vibe code Vision Pro apps with Siri

Huawei Enjoy 80 appears in live photos

AMD’s Ryzen 7 7800X3D hardware bundle at Micro Center is $80 off

Discover more from rjema