The assessment revealed a widening gap between top performers and stragglers, including xAI, Meta $META, and Chinese companies DeepSeek, Z.ai, and Alibaba Cloud. Companies across the board performed poorly in the Current Harms domain, which evaluates how AI models perform on standardized trustworthiness benchmarks that test designed to measure safety, robustness, and the ability to control harmful outputs. Reviewers found that "frequent safety failures, weak robustness, and inadequate control of serious harms are universal patterns" with uniformly low performance on these benchmarks.