01 · the premise
Build voice and vision products people love.
Understanding is the new interface.
Does it catch the tone, the timing, the glance and the gesture, or just the words? xevall is the independent evals and self-improvement layer for human-AI interaction.
- processing · on this device
- media transmitted · 0
- system graded · never the person
02 · natural inputs
Interfaces are losing their keyboards.
What replaces typing is you: speech, tone, gaze, gesture. Products built around natural input enter the examiner's field. Their understanding is what gets tested.
- gaze
- blink
- speech
- tone
- gesture
03 · system misfires
Nobody measures the understanding.
When an AI system misunderstands a person, the product feels broken and the reason stays hidden. xevall grades the system against the meaning a human reviewer established, and hands back the moment it went wrong.
| session | human meaning | system action | grade |
|---|---|---|---|
| kiosk 04:12 | a pause, not a yes | took it as a yes | misfire · consent |
| agent 00:31 | still mid-thought | jumped in | misfire · timing |
| telehealth 11:47 | said fine, meant not fine | noticed, and asked | understood · 0.92 |
- intended
- inferred
05 · the improvement loop
Every interaction creates traces to improve your product.
You already eval the words. Nobody is evaluating the coherence between the modalities: the tone, the distance, the body language. xevall grades whole interactions against human judgement, marks where the understanding broke, and re-scores the same sessions once you have fixed it. The grading exists to make products more human, not to mint a number.
- agreement 0..1
the Human Input Benchmark
The Human Input Benchmark goes public this autumn.
Building a product around voice, timing, gaze or gesture? Early access is open for teams testing natural interfaces.
xevall ·the independent evals and self‑improvement layer for human‑AI interaction.