Measured, not marketed
No mind-reading — measurable acoustic and visual signal, reported with its dataset, sample size, and honest caveats. Where a signal is weak or unproven, we label it plainly.
99.7% recall · 99.3% specificity
Held-out on 979 real breed barks + 300 ESC-50 non-bark negatives. Small field set (n=34): 100% / 100%. This is the part we trust most.
R² ≈ 0.19 · 40% vs 34% chance
Supervised head, held-out Barkopedia n=192 (husky/shiba scope). Weak-but-real and validated — it leads the emotion read. Breed-stratified it's uneven (shiba R² 0.27, husky 0.12), so it's not yet proven across breeds — broadening is data-gated.
At chance
Every audio model we tried fails to beat chance on valence — a bark encodes energy, not positivity. We say so, and we don't oversell it.
AUC 0.930 · 84% accuracy
A CLIP vision head reads body-language valence (held-out on the Dewa dog-emotion set). Multi-frame video 0.907; robust to real webcam capture (−0.006). This is how we solve the axis audio can't.
≈1% false-alarm rate
Rather than an absolute emotion, we track how a dog's own voice moves against its own 3-week robust baseline (median/MAD). An off-day needs several acoustic features to shift together, tuned for a low false-positive rate — measured at ≈1% per day on stable synthetic data (target 5%). A prompt to look, never a diagnosis.
33% top-1 (chance 20%)
Better than chance, but deliberately for entertainment — shown as a playful spread, never a confident claim.