a16z just dropped $40 million into Vals AI. A Series A for an AI evaluation tool. The narrative is clean: reliable AI assessment is critical. The market needs it. The press parrots the line.
But I have seen this playbook before. In 2020, I spent three months verifying zk-Rollup proofs. The whitepaper promised trustless scalability. The code revealed a fraud proof window discrepancy. The market believed the narrative. I found the flaw.
Vals AI's announcement is a snapshot. Not a guarantee. The technical details are absent. The product is a black box. The funding is a bet on a category, not on a verifiable technology.
Let me break down what we actually know.
Hook: The $40M Signal
$40 million is a large Series A. The implied valuation is likely $150M to $200M. That is a strong signal. But signals are not evidence. The signal says: a16z believes AI evaluation is the next infrastructure layer. The signal does not say: Vals AI has solved the fundamental problems.
The article from Crypto Briefing is thin. Five information points. No technical white paper. No open-source code. No comparison with existing tools like LangSmith, Galileo, or Patronus AI. This is a common pattern in bull markets. Money flows into narratives. Technical rigor is an afterthought.
Context: The AI Evaluation Stack
AI evaluation tools sit between model development and production deployment. They measure accuracy, safety, robustness. They are the software QA for AI. In 2024-2025, the industry shifted from static benchmarks (MMLU, HellaSwag) to agentic evaluation. Complex multi-step tasks. Autonomous decision-making.
This is exactly where crypto meets AI. Smart contracts that call LLMs. AI agents that execute trades. DeFi protocols that rely on model outputs. These systems need evaluation frameworks that go beyond academic benchmarks. They need adversarial testing. They need formal verification.

I know this because in 2025, I designed a formal verification framework for AI-agent smart contract interactions. The core problem is not the model. It is the evaluation methodology. How do you prove that an agent will not execute a malicious transaction? How do you measure the risk of prompt injection?
Vals AI claims to address this need. But the article does not tell us how.
Core: What Vals AI Probably Does — And What It Hides
Based on the company's positioning, Vals AI is an evaluation tool provider. It likely uses a "LLM-as-Judge" architecture. This means it calls models like GPT-4o, Claude, or Gemini to evaluate outputs of other models. The judge is another AI.
This is common. It is also a vulnerability. The judge itself is fallible. The evaluation is only as good as the judge's calibration. The field calls this the "evaluator evaluation problem" — who audits the auditor?
Vals AI's tool probably includes:
- Custom evaluation datasets
- Scenario-based test generation
- Automated scoring with rubric
- Dashboard for tracking model performance over time
But the article does not specify evaluation dimensions. Does it measure robustness against adversarial attacks? Does it detect hallucinations in financial contexts? Does it support multi-modal inputs? Does it handle agentic workflows where the model executes multiple steps?
These are the technical differentiators. Without them, Vals AI is indistinguishable from a dozen other startups.
Check the math, not the roadmap. The investment is based on a roadmap. The market needs to check the math.
Contrarian: The Audit Theater Trap
Here is the contrarian angle. Evaluation tools can create a false sense of security. Companies run a benchmark. The score is high. They deploy the model. The real-world failure rate is still high. Why? Because the evaluation did not cover the edge cases.
This is "audit theater" — the appearance of scrutiny without the substance. I have seen this in blockchain audits. A smart contract passes a security audit. The audit is a snapshot. The code changes. The threat landscape changes. The snapshot expires.
Audits are snapshots, not guarantees. The same applies to AI evaluation. Vals AI's tool is a snapshot of a model's performance on a specific dataset. It does not guarantee performance in production.
Moreover, there is a deeper problem. Complexity is the enemy of security. AI evaluation tools are themselves complex software. They have bugs. They have biases. They can be gamed. If Vals AI's tool is a black box, its users cannot verify its accuracy. They are trusting a trust machine.
In my experience building formal verification for AI agents, the only way to ensure reliability is to make the evaluation framework open-source and transparent. The datasets must be public. The scoring algorithm must be auditable. Otherwise, the tool is a vendor lock-in, not a safety net.
Vals AI has not disclosed any of this. The article does not mention open-source plans. It does not mention third-party audits of the evaluation tool itself. This is a red flag.
Takeaway: The Burden of Proof
$40 million is a bet on the thesis that AI evaluation is essential. The thesis is correct. The industry needs robust, independent evaluation tools. But the execution is unproven.

Until Vals AI publishes its methodology, its code, its benchmark results, and its adversarial test suite, the investment is speculation. The market is euphoric. The funding is a signal of FOMO, not of technical validation.
Code does not care about your vision. Vals AI's vision is compelling. The code is unknown. The industry should demand proof before assuming trust.
I will be watching. If Vals AI releases a transparent evaluation framework, I will analyze it. If it remains a black box, I will treat it as another product looking for a problem.
The AI evaluation race is just beginning. The winners will be those who prioritize verifiability over narrative. The losers will be those who assume that a $40M check equals a technical breakthrough.