OpenCV AI Competition 2026

Your camera,
your coach.

Doovio watches you practise through your phone camera, counts and checks what you do, tells you exactly what to fix, and says so when it can't see well enough to judge.

Doovio coaching the ASL letter B: hand skeleton, the target letter card and the correction 'That looks like V. Straighten your little finger.'
Doovio"That looks like V. Straighten your little finger."

Four skills, one coach

Each skill has its own way of seeing and its own lessons. They all share the same honest rule: no feedback unless the camera actually saw it.

Squat with skeleton overlay and rep counter

Gym

Counts squats and push-ups, checks depth, chest position and hip line from a side view, and adapts to your range once you approve it.

2 lessons
Assembly board with the target square highlighted and a correction

Assembly

Coloured blocks on a printed board. Every step is verified before the next one, and mistakes are named precisely: "the blue block is on the top-middle square".

2 lessons
Sign language lesson with hand skeleton and letter card

Sign Language

ASL fingerspelling, letter by letter, with a picture of each handshape and corrections like "bend your ring finger". A practice aid, not a replacement for Deaf teachers.

4 lessons
Presenting lesson: face box and the nudge 'Look at the camera'

Presenting

A practice talk where the camera is your audience (eye contact, hands, swaying) and a guided calm-breathing exercise checked from your chest movement.

2 lessons

Coming next: Piano (camera plus microphone or MIDI), Cooking, Drawing, Driving practice. They are shown as "coming soon" in the app and cannot give feedback until they really work.

How it works

A few frames per second go from the phone to the Doovio server, which sees, decides and answers within tens of milliseconds.

1
Can I see?Brightness, sharpness and camera shake. If not, "Cannot assess — adjust the camera".
2
What is there?OpenCV 5: body, hand and face models via cv.dnn, ArUco boards, optical flow, Kalman smoothing.
3
Is it right?Rules over recent frames: rep counting, step verification, letter matching, habits.
4
Say one thing.A single specific correction or a confirmation, spoken aloud, never repeated every second.
Architecture diagram: phone and web clients, OpenCV 5 pipeline and vision agent on an AWS Graviton instance, ECR, Secrets Manager, CloudWatch, Bedrock
Runs on an AWS Graviton (arm64) instance: image in ECR, key in Secrets Manager, logs and per-session metrics in CloudWatch, agent planning on Amazon Bedrock.
Agentic vision

When the coach is stuck, an agent takes a closer look

Claude on Amazon Bedrock is called only when the rule-based coach gets stuck. It measures the scene with OpenCV tools, decides, acts, and checks the result. It never sees an image.

  • Tools return numbers: which markers are visible, joint angles, finger bends
  • Any change to how you are coached needs your "Yes"
  • Works offline too, with a rule planner using the same tools
  • Every step is logged in the session trace
11.7s PERCEPTION squares on the right edge unreadable for 5 s
11.7s AGENT look_at_camera → corners not visible: top-right, bottom-right · move_phone_toward: right
14.6s ACTION "Move the phone to the right so all corners of the board are in view."
18.7s AGENT wait_and_look → all 4 markers visible
19.0s DECISION step 2 verified

Honest results

Measured with the real pipeline on labelled videos and photos, including clips we never tuned on. Failures are reported too.

0
false reps on wrong-angle, dark or blurred video
in 41 runs; 2 in 32 more runs on a fresh held-out set, all on one heavily cropped clip.
4/4
assembly steps verified, 2/2 mistakes caught
on rendered test clips at 3 to 25 fps, with no false mistake reports and no corrections while the board was out of view.
91%
ASL letters read correctly (dataset test)
70% on a new signer in a live test. Look-alike fist letters such as A and S are the weak spot.
8/9
right agent actions (Claude, 3 attempts per scenario)
15/15 right triggers; never acted when it wasn't needed.
22–34 ms
per frame on AWS Graviton
median server time for sign and gym frames on a t4g.small instance.
5/5
presenter clips judged correctly
eye contact, hands out of view and looking away; breathing logic is still to be validated on people following the count.

These are small test sets. They show how the system behaves, not statistically robust accuracy. Details, failure cases and limits are in the technical report.

Private by design

A coach that watches you has to earn trust.

Nothing is recordedFrames are analysed in memory and discarded. Sessions are deleted when they end.
Asked firstThe app asks for consent before the camera is used for the first time.
No mind readingDoovio coaches visible behaviour. It does not detect emotions or stress from your face, and gym feedback is not medical advice.

See it in action

A five-minute walkthrough: the team, the app working, the architecture and the results.

Demo video coming soon.
Meanwhile, try the live demo: sample videos run through the real pipeline.

Try Doovio now

Use the sample videos or your own webcam. No install, no account.

Open the live demo