Kunko is building the AI Decision Assurance Platform for enterprises. We help companies audit, control and continuously assure AI systems before allowing them to make autonomous decisions.
We start with open-source Judge Audit, which measures whether AI judges can actually be trusted. We are expanding this into a full assurance platform covering AI decisions, agents, evidence, policies and continuous monitoring.
Not another governance dashboard. Kunko is the layer that continuously determines what AI is allowed to do — and whether it remains trustworthy enough to do it. Three planes, one evidence backbone:
Monitor → detect change → re-evaluate → act. When the model, prompt, data or policy changes, Kunko re-runs the evidence and can automatically downgrade or block autonomy. Trust is not a one-time certificate — it's a stream.
Evaluations · calibration reports · audit trails · provenance · signed dossiers — everything the platform decides is backed by evidence a regulator, an auditor or your board can inspect.
What AI systems do I run? Model, agent, judge, prompt, version, owner, environment, data, provider.
What risk does each carry? Low, medium, high, critical — per system, per decision type.
Why do we trust it? Evaluations, calibration, red-team results, security tests, human outcomes, incidents, audit history.
What may it do? Automate · review · block — per decision, per risk level.
Judge Audit, our open-source engine, runs AI judges in shadow mode against decisions humans already made. No rip-and-replace, no blind trust — just measurement:
Safe Automation Rate turns reliability into economics: how much human work can you safely remove? The value we create is directly proportional to the amount of human review — and risk — we safely remove.
pip install kunko-judge-audit · github.com/kunko-ai-labs/judge-audit →
Our v0.5 study ran three judges over 3,080 human-labelled banking decisions (BANKING77) — pre-registered before the first model call, every result published including the unflattering ones. Same data, same 5% error target:
| Judge | Decides alone at ≤ 5% error | Range over seeds |
|---|---|---|
| TypeSafe Jev, native probability | 41.0% | 27.0 – 45.1% |
| Gemini 3.6 Flash, verbalized | 14.3% | 7.6 – 30.0% |
| Qwen3-8B, token log-probability | 0% | 0.0 – 4.4% |
A 95% confidence score is not a 95% guarantee — token probabilities couldn't automate a single decision at this bar. That's the point: confidence means different things across judges, and Kunko measures whether it can be trusted. Caveats as published: public dataset, likely seen in pretraining; label noise not measured. Full study → The v0.5 roster is the start — Judge Arena keeps score as it grows.
From AI evaluation to AI assurance. The market is full of "does my model work?". We take you to "can I safely trust this system with this decision?" — and then to "can I continuously prove it remains safe to trust?"
A bank processes 200,000 customer-service decisions a month. Today 100% require human review. Kunko proves that 63% can be automated while keeping error below 2% — and continuously monitors whether that remains true.
Or an insurer reviewing claims: which auto-approve, which need a human, which must be blocked. Or a fintech running AML triage inside its risk threshold. One workflow, one measured number, one budget line that shrinks.
Give us 5,000 of your historical decisions. We'll show you exactly what percentage of your current human review could be automated — under your own error threshold.
Start in shadow mode. No production risk.
Head of AI / AI Platform / ML Engineering — owns the models making the decisions.
COO / Operations / Risk / Transformation — owns the cost of human review.
CISO / Model Risk / Compliance — owns what happens when the AI is wrong.
Our initial ICP: AI teams in financial services and other regulated enterprises deploying AI into high-volume decisions that still require human review — because nobody can prove when the AI is safe to trust.
The measurement layer, open to everyone. We open-sourced it because the industry needs an independent standard for AI decision reliability — and because credibility compounds: every public audit makes Kunko the authority on what "trustworthy" means.
The commercial product is not the auditor — the auditor is the entry point. We monetize the assurance layer around enterprise AI decisions: continuous monitoring, evidence dossiers, policies, risk controls, integrations and governance.
We're deliberately starting with an open-source wedge: Judge Audit builds credibility and adoption around an independent standard before asking enterprises to put sensitive production decisions behind a commercial platform. Our job is to remain model-independent — the company whose model benefits from being trusted shouldn't be the only one certifying that trust.
Judge Audit isn't the company — it's the wedge. Each layer compounds on the last:
Open-source calibration audits. The credibility layer — live today.
Evidence + risk + policy per decision: automate, review, or block.
Capability control for agents: what it may do, verified continuously.
The full control plane: inventory, risk, evidence, policy, monitoring.
Continuous assurance: when anything changes, Kunko re-evaluates the evidence and downgrades or blocks autonomy automatically.
The vendor tells you how confident its model is. We independently measure whether that confidence is trustworthy for your decisions — and we stay put when you swap the model.
Every audited decision generates evidence: decision, model, prompt, confidence, ground truth, risk, outcome, override. Over time, that's proprietary data on when enterprises actually trust AI — and when it fails them.
Safe Automation Rate converts reliability into budget: the share of human review you can remove at a certified error. Nobody else sells that number.
Judges decide; agents act — the second building block of the platform. Declare what your AI agent may do, and verify it on every edit, PR and release. A GitHub Action that checks MCP configs, Claude Code permissions and tool definitions against your policy — deterministic, OWASP-mapped, with signed evidence.
"We don't want to make AI more confident.
We want to make autonomy measurable."
As AI moves from generating information to making decisions and taking actions, enterprises need a system that continuously determines what AI is allowed to do — and whether it remains trustworthy enough to do it. That's Kunko.
If an AI judge says it's 90% confident across 100 decisions, it should be right about 90 times. If it's right 60 times, it's miscalibrated — overconfident. Calibration is the match between stated confidence and actual accuracy, and it's what makes a judge safe to automate around.
No. Accuracy tells you how often the judge is right; calibration tells you whether you can believe it when it matters. A 95%-accurate judge reporting 99.9% confidence on its errors is more dangerous than an 80%-accurate judge that knows when it's unsure. Automation decisions need the second number.
The vendor can tell you how confident its model is. It cannot independently tell you whether that confidence is trustworthy for your business decisions. Kunko sits outside the model: we measure the decision system against your own ground truth, risk thresholds and policies — and we stay put when you swap the underlying model.
No. Judge Audit judges nothing — it measures. Given a judge and a dataset with known-correct answers, it computes statistics. That's arithmetic, not opinion, and every number is reproducible from raw checkpoints in the repo.
High-risk AI systems (in force December 2027) need logging, human oversight, and demonstrated accuracy and robustness. Independent calibration evidence — including worst-case metrics, not just averages — is exactly the kind of documentation those obligations call for.
pip install kunko-judge-audit, bring your judge and a labeled dataset, and follow the quickstart in the repo. Pre-register your hypotheses if you want the audit to count as evidence.
Audits, partnerships, enterprise pilots — or 5,000 historical decisions and an error threshold.
[email protected]