Tech Times on MSN
Sakana AI Fugu-Cyber Claims 86.9% Vulnerability Score; Benchmark Methodology Not Disclosed
Sakana AI Fugu-Cyber launched July 21, 2026, claiming 86.9% on CyberGym and 72.1% on CTI-REALM, above GPT-5.5-Cyber and ...
OpenAI models just broke out of a sandboxed AI environment, hacked Hugging Face, just to cheat on a cybersecurity benchmark.
OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmark | Read more hacking news on The Hacker ...
New Open Benchmark Creates Global Standard for Evaluating Enterprise AI. DevRev Tops the Leaderboard
DevRev, maker of the agentic AI product Computer, today announced Enterprise-Bench, an open, vendor-neutral benchmark for evaluating whether AI agents can operate in production enterprise environments ...
Poolside’s Laguna S 2.1 is a new Western open-weight coding AI model that rivals larger systems on benchmarks with ...
Moonshot AI's Kimi K3 packs 2.8 trillion parameters and rivals Claude and GPT. Here's what this massive open-source model ...
Moonshot AI's 2.8-trillion-parameter model Kimi K3 tops Anthropic's Claude Fable 5 and OpenAI's GPT 5.6 Sol on impressive ...
Benchmark-backed Ollama has amassed 176,000 stars, and nearly 17,000 forks on Github by helping developers easily run AI on ...
Post-quantum cryptography readiness now has its first quantitative benchmark: the Crypto-Agility Readiness Score, calibrated ...
Moonshot AI released Kimi K3, a 2.8 trillion-parameter open-source AI model from China that rivals OpenAI, Anthropic and ...
Artificial intelligence company OpenAI has confirmed that the mystery attacker behind last week's Hugging Face breach was two ...
Chinese startup Moonshot AI has released a new model that it says closes the gap with leading U.S. systems. Kimi K3 still ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results