Models From Score Interview

OpenAI o3 Model: Lower Benchmark Scores Raise Questions About Claims, Transparency Over AI

OpenAI has long been touting the capabilities of its artificial intelligence (AI) developments, especially with their o-series models that are capable of reasoning and more advanced capabilities. The ...

TechCrunch

OpenAI’s o3 AI model scores lower on a benchmark than the company initially implied

A discrepancy between first- and third-party benchmark results for OpenAI’s o3 AI model is raising questions about the company’s transparency and model testing practices. When OpenAI unveiled o3 in ...

TechCrunch

One of Google’s recent Gemini AI models scores worse on safety

A recently released Google AI model scores worse on certain safety tests than its predecessor, according to the company’s internal benchmarking. In a technical report published this week, Google ...

Some results have been hidden because they may be inaccessible to you

Show inaccessible results

OpenAI o3 Model: Lower Benchmark Scores Raise Questions About Claims, Transparency Over AI

OpenAI’s o3 AI model scores lower on a benchmark than the company initially implied

One of Google’s recent Gemini AI models scores worse on safety

Trending now