CollectivIQ Achieves Unparalleled Accuracy in AI
In a groundbreaking announcement, Boston-based CollectivIQ has demonstrated remarkable advancements in the field of artificial intelligence (AI) by outperforming well-known frontier models in accuracy, achieving an impressive 96.4% GPQA Diamond accuracy score. As the world's first AI consensus platform for business intelligence, this development underscores the pivotal role of collaborative intelligence in elevating AI performance.
The company's proprietary consensus algorithm facilitates queries across numerous leading large language models (LLMs) concurrently—such as ChatGPT, Gemini, Claude, and Grok—allowing it to synthesize information effectively. By identifying commonalities and disparities among different responses, CollectivIQ can deliver answers that are not only accurate but also grounded in reliability and integrity.
Delivering Higher Quality Intelligence
CollectivIQ's exceptional performance was verified through independent evaluations conducted by Ten Point Data, revealing that it achieved a striking 53.3% on the Humanity’s Last Exam (text-only), securing the highest score among its competitors. Its GPQA Diamond accuracy leads globally, exceeding the typical human PhD baseline by a substantial 26 percentage points. Notably, CollectivIQ managed to do this with earlier versions of frontier models, signaling its robustness even when not utilizing the latest technology.
Another critical focus of the assessment was calibration— the ability of a model to predict its uncertainty. CollectivIQ recorded an Expected Calibration Error (ECE) of 0.41, which represents a remarkable improvement of 28.1% compared to traditional single model architectures, which average an ECE of 0.57. This is particularly crucial in high-stakes enterprise environments where the costs of inaccurate information can be extreme.
A Transformative Approach to AI
John Davie, CEO of CollectivIQ, expressed his enthusiasm, stating, "The age-old saying 'two heads are better than one' is being validated in the field of AI. By merging insights from various advanced AI models through our consensus engine, CollectivIQ offers superior intelligence that is both economical and less prone to the inaccuracies that often plague singular approaches."
As businesses grapple with the financial impact of AI misfires, which are estimated to cost global companies around $67.4 billion annually, the need for dependable AI solutions becomes evident. Current trends show that employees now spend over 51 workdays per year contending with technological disruptions, which has surged 42% since 2025. Significant time is dedicated to validating AI outputs, further underscoring the significance of solutions that reduce friction and enhance accuracy.
Optimizing AI for Enterprises
The design of CollectivIQ as an asynchronous expert employee allows it to leverage the most suitable LLMs for various tasks while balancing cost and intelligence levels. Routine inquiries can be addressed efficiently through cost-effective models, while complex issues are directed to high-performing, robust models. This prioritization of factual accuracy over rapid response times is vital in enterprise settings, where incorrect answers can be detrimental.
Founded by a team with deep roots in the AI domain, CollectivIQ is revolutionizing access to AI technology by merging leading LLMs into a coherent and accountable framework. Its objective is to enhance governance and economic alignment within enterprise AI strategies. For more information on how CollectivIQ is setting new standards in the AI landscape, visit
www.collectiviq.ai.