About

What this is

Komplex AI builds reliability infrastructure for large language models. Our first product is a hallucination detector that takes an LLM response (and optionally the prompt) and returns a calibrated probability that the response contains a hallucination.

What it isn't: It's not a judge model. It doesn't re-generate the answer to compare versions. It doesn't hit a retrieval index, search the internet, or cross-check against an external fact database. A single fast pass — typically sub-second on a warm path — with no infrastructure on your end.

Probabilistic, not perfect. The detector is a triage signal, not a definitive verdict, and some hallucination types are caught much more reliably than others. See the Performance page for current accuracy numbers.

Why it matters

LLMs hallucinate. The problem is not going away — it is a structural property of how these models work. Every team deploying an LLM in production needs a way to flag unreliable outputs before they reach users.

We make that easy: one API call, one probability score, no infrastructure required.

How it works

Send a prompt and a response (or just a response); receive a probability score and per-hallucination-type breakdown. See Performance for accuracy by hallucination type.

Contact

Questions, feedback, or partnership inquiries: get in touch.