勲功  ·  Kunko — distinguished merit

Kunko tells enterprises
which AI decisions are safe to automate.

Kunko is building the AI Decision Assurance Platform for enterprises. We help companies audit, control and continuously assure AI systems before allowing them to make autonomous decisions.

We start with open-source Judge Audit, which measures whether AI judges can actually be trusted. We are expanding this into a full assurance platform covering AI decisions, agents, evidence, policies and continuous monitoring.

Explore Judge Audit Request the Enterprise Demo
The vision

The control plane for AI assurance

Not another governance dashboard. Kunko is the layer that continuously determines what AI is allowed to do — and whether it remains trustworthy enough to do it. Three planes, one evidence backbone:

DISCOVER
  • AI inventory — every system you run
  • Models, agents, judges, prompts
  • Versions, owners, environments
  • Data sources & providers
ASSURE
  • Judge Audit — calibration evidence
  • Agent Assurance — capability control
  • Evaluations, red-team, security tests
  • Human outcomes & incident history
CONTROL
  • Risk engine — low → critical per system
  • Policies — what each system may do
  • Human review queues
  • Runtime controls & blocks
CONTINUOUS ASSURANCE

Monitor → detect change → re-evaluate → act. When the model, prompt, data or policy changes, Kunko re-runs the evidence and can automatically downgrade or block autonomy. Trust is not a one-time certificate — it's a stream.

EVIDENCE LAYER

Evaluations · calibration reports · audit trails · provenance · signed dossiers — everything the platform decides is backed by evidence a regulator, an auditor or your board can inspect.

Inventory

What AI systems do I run? Model, agent, judge, prompt, version, owner, environment, data, provider.

Risk

What risk does each carry? Low, medium, high, critical — per system, per decision type.

Evidence

Why do we trust it? Evaluations, calibration, red-team results, security tests, human outcomes, incidents, audit history.

Policy

What may it do? Automate · review · block — per decision, per risk level.

How it works

From AI decision to controlled autonomy

Judge Audit, our open-source engine, runs AI judges in shadow mode against decisions humans already made. No rip-and-replace, no blind trust — just measurement:

AI decidesjudge or agent output
→
Confidence?what it claims
→
Trusted?calibration check
→
Measured errorcertified bounds
→
% automatablesafe automation rate
→
Risk OK?your threshold
AUTOMATE REVIEW BLOCK

Safe Automation Rate turns reliability into economics: how much human work can you safely remove? The value we create is directly proportional to the amount of human review — and risk — we safely remove.

pip install kunko-judge-audit  ·  github.com/kunko-ai-labs/judge-audit →

Proof, not promises

The same audit, three different truths

Our v0.5 study ran three judges over 3,080 human-labelled banking decisions (BANKING77) — pre-registered before the first model call, every result published including the unflattering ones. Same data, same 5% error target:

JudgeDecides alone at ≤ 5% errorRange over seeds
TypeSafe Jev, native probability41.0%27.0 – 45.1%
Gemini 3.6 Flash, verbalized14.3%7.6 – 30.0%
Qwen3-8B, token log-probability0%0.0 – 4.4%

A 95% confidence score is not a 95% guarantee — token probabilities couldn't automate a single decision at this bar. That's the point: confidence means different things across judges, and Kunko measures whether it can be trusted. Caveats as published: public dataset, likely seen in pretraining; label noise not measured. Full study → The v0.5 roster is the start — Judge Arena keeps score as it grows.

Evaluationdoes my model work?
→
Assurancecan I trust this decision?
→
Autonomyautomate / review / block
→
Continuous assuranceprove it stays safe

From AI evaluation to AI assurance. The market is full of "does my model work?". We take you to "can I safely trust this system with this decision?" — and then to "can I continuously prove it remains safe to trust?"

Under the hood

We measure what accuracy can't tell you

stated confidence → actual accuracy perfect this judge
A reliability diagram: the diagonal is perfect calibration — most judges aren't on it.
  • Safe Automation RateWhat share of decisions can be automated at a certified error bound. The ROI number.
  • ECE with confidence intervalsHow far confidence drifts from reality — as an interval, never a lonely point estimate.
  • MCE — worst-bin errorThe regulator's question, not the average case: where is this judge most wrong?
  • Selective prediction, certifiedFinite-sample risk bounds (Clopper–Pearson): automation you can sign off on.
  • Adversarial robustnessPrompt injection, homoglyphs, social engineering — does calibration survive attack?
  • Cost, latency & driftCalibration per dollar and per millisecond — plus continuous monitoring for when it degrades.
The killer use case

From 100% human review to measured autonomy

A bank processes 200,000 customer-service decisions a month. Today 100% require human review. Kunko proves that 63% can be automated while keeping error below 2% — and continuously monitors whether that remains true.

Or an insurer reviewing claims: which auto-approve, which need a human, which must be blocked. Or a fintech running AML triage inside its risk threshold. One workflow, one measured number, one budget line that shrinks.

Kunko Enterprise Assurance Demo

Give us 5,000 of your historical decisions. We'll show you exactly what percentage of your current human review could be automated — under your own error threshold.

Start in shadow mode. No production risk.

Champion

Head of AI / AI Platform / ML Engineering — owns the models making the decisions.

Economic buyer

COO / Operations / Risk / Transformation — owns the cost of human review.

Security buyer

CISO / Model Risk / Compliance — owns what happens when the AI is wrong.

Our initial ICP: AI teams in financial services and other regulated enterprises deploying AI into high-volume decisions that still require human review — because nobody can prove when the AI is safe to trust.

Business model

Open source is the wedge.
The platform is the business.

FREE · OPEN SOURCE

Judge Audit

The measurement layer, open to everyone. We open-sourced it because the industry needs an independent standard for AI decision reliability — and because credibility compounds: every public audit makes Kunko the authority on what "trustworthy" means.

ENTERPRISE

Kunko Assurance Platform

The commercial product is not the auditor — the auditor is the entry point. We monetize the assurance layer around enterprise AI decisions: continuous monitoring, evidence dossiers, policies, risk controls, integrations and governance.

We're deliberately starting with an open-source wedge: Judge Audit builds credibility and adoption around an independent standard before asking enterprises to put sensitive production decisions behind a commercial platform. Our job is to remain model-independent — the company whose model benefits from being trusted shouldn't be the only one certifying that trust.

Roadmap

Judge Audit is the entry point
into a much bigger control plane.

Judge Audit isn't the company — it's the wedge. Each layer compounds on the last:

Judge Audit

Open-source calibration audits. The credibility layer — live today.

Decision Assurance

Evidence + risk + policy per decision: automate, review, or block.

Agent Assurance

Capability control for agents: what it may do, verified continuously.

Enterprise Assurance

The full control plane: inventory, risk, evidence, policy, monitoring.

CERTIFIED→ model / prompt / data changes→ IMPACT ANALYSIS→ REVALIDATE→ CERTIFIEDDEGRADEDBLOCKED

Continuous assurance: when anything changes, Kunko re-evaluates the evidence and downgrades or blocks autonomy automatically.

Why Kunko wins

We don't decide whether AI is intelligent.
We determine whether it's trustworthy enough to act.

Independence

Outside the model

The vendor tells you how confident its model is. We independently measure whether that confidence is trustworthy for your decisions — and we stay put when you swap the model.

Evidence moat

Compounding ground truth

Every audited decision generates evidence: decision, model, prompt, confidence, ground truth, risk, outcome, override. Over time, that's proprietary data on when enterprises actually trust AI — and when it fails them.

Economics

A number, not a dashboard

Safe Automation Rate converts reliability into budget: the share of human review you can remove at a certified error. Nobody else sells that number.

Also from the lab

Agent Assurance

Judges decide; agents act — the second building block of the platform. Declare what your AI agent may do, and verify it on every edit, PR and release. A GitHub Action that checks MCP configs, Claude Code permissions and tool definitions against your policy — deterministic, OWASP-mapped, with signed evidence.

github.com/kunko-ai-labs/agent-assurance →

"We don't want to make AI more confident.
We want to make autonomy measurable." As AI moves from generating information to making decisions and taking actions, enterprises need a system that continuously determines what AI is allowed to do — and whether it remains trustworthy enough to do it. That's Kunko.

FAQ

Questions we get

What does "calibration" mean, in plain language?

If an AI judge says it's 90% confident across 100 decisions, it should be right about 90 times. If it's right 60 times, it's miscalibrated — overconfident. Calibration is the match between stated confidence and actual accuracy, and it's what makes a judge safe to automate around.

Isn't accuracy enough?

No. Accuracy tells you how often the judge is right; calibration tells you whether you can believe it when it matters. A 95%-accurate judge reporting 99.9% confidence on its errors is more dangerous than an 80%-accurate judge that knows when it's unsure. Automation decisions need the second number.

Why can't the model vendor do this?

The vendor can tell you how confident its model is. It cannot independently tell you whether that confidence is trustworthy for your business decisions. Kunko sits outside the model: we measure the decision system against your own ground truth, risk thresholds and policies — and we stay put when you swap the underlying model.

Is Judge Audit itself another AI judge?

No. Judge Audit judges nothing — it measures. Given a judge and a dataset with known-correct answers, it computes statistics. That's arithmetic, not opinion, and every number is reproducible from raw checkpoints in the repo.

Why does this matter for the EU AI Act?

High-risk AI systems (in force December 2027) need logging, human oversight, and demonstrated accuracy and robustness. Independent calibration evidence — including worst-case metrics, not just averages — is exactly the kind of documentation those obligations call for.

How do I run an audit?

pip install kunko-judge-audit, bring your judge and a labeled dataset, and follow the quickstart in the repo. Pre-register your hypotheses if you want the audit to count as evidence.

Contact

Let's measure what matters

Audits, partnerships, enterprise pilots — or 5,000 historical decisions and an error threshold.

[email protected]