Kamiwaza AI, the secure AI orchestration platform that helps enterprises turn distributed data into secure, governed AI action, today announced that Signal65 PINNACLE, a new independent benchmarking framework, co-developed with and powered by Kamiwaza technology, is generally available for enterprise evaluations. This new framework is designed to measure whether AI systems successfully complete enterprise work, how much correct work infrastructure can sustain, and what each correct outcome costs.
At the foundation of this approach is technology developed by Kamiwaza, including its PICARD execution harness and KAMI and RIKER benchmark work. PICARD enables agents to perform multi-step jobs against generated enterprise environments using tools such as filesystems, databases and Python runtimes. Each environment is created with a known answer key held outside the agent's reach, allowing the completed work to be objectively graded in code rather than evaluated by another AI model or human reviewer.
"Enterprises are being asked to make major AI infrastructure decisions on benchmarks that tell them the wrong story," said JV Roig, Senior Generative AI Platform Engineer, developer of the technology powering PINNACLE, Kamiwaza AI. "Much of what these tests contain is already sitting in the training data; it's impossible to separate memorization vs real capability. Even a clean score doesn't measure the capability enterprises need including multi-step jobs based on occupational skills, with real tools, messy data, decision-making under ambiguity, and resisting the tendency to hallucinate. Plus, it isn't only about the model, it's whether the combination you've chosen can actually complete the work your business needs done, reliably, at the scale you require and at a cost that makes sense. We built this technology to make that measurable, and Signal65's use of it as the foundation for PINNACLE shows how much that capability matters now."
The first PINNACLE data set validates why enterprises cannot identify the right AI model or infrastructure configuration through conventional benchmarks alone and instead need workload-specific evaluation. Signal65 tested 44 configurations across 30 base models from 12 makers, including hosted APIs and open-weight models running on NVIDIA H200, NVIDIA B300 and AMD Instinct MI300X infrastructure. Among the initial findings:
- Performance varied significantly by data conditions, as nine configurations completed at least 95% of jobs with clean, governed data, but only two reached that threshold with messy data.
- Rankings changed by workload, with smaller and open-weight models competing with or outperforming frontier models on certain tasks.
- Token pricing did not reflect the cost of successful work, with input tokens representing 65% to 91% of hosted-model task costs, while output-token pricing alone understated the cost per correct task by three to 11 times.
The findings reinforce a core principle behind Kamiwaza's approach to AI orchestration: there is no single model or infrastructure configuration that is right for every enterprise workload. Kamiwaza's model- and infrastructure-agnostic orchestration gives enterprises the flexibility to make those choices based on what works best for their specific environment rather than being locked into a single stack.
"These results validate why we've built Kamiwaza around orchestration and choice from the beginning," said Luke Norris, CEO and co-founder of Kamiwaza AI. "Enterprises shouldn't have to bet their AI strategy on one model or infrastructure provider, hoping it works across every workload. I'm incredibly proud of the pioneering work our team continues to develop to help enterprises not only orchestrate across models, infrastructure, data and agent, but make those decisions based on what truly delivers optimum outcomes for their business."
"Agentic AI changes what enterprises need from benchmarking," said Ryan Shrout, President and General Manager of Signal65. "It isn't enough to know how capable a model is in isolation or how quickly a piece of hardware can generate tokens. Enterprises need to know whether the entire system can complete useful work correctly and economically. Kamiwaza's technology gave us the foundation to build a benchmark around that question, and Signal65 PINNACLE gives enterprises a way to start answering it with objective data."