Arduino

CantaStorie: Offline AI Storytelling Toy on Arduino UNO Q

CantaStorie: Offline AI Storytelling Toy on Arduino UNO Q

Parts list first: an Arduino UNO Q, a small camera, a speaker, an addressable LED ring, two potentiometers, and a battery pack. That is roughly everything inside CantaStorie, a screen-free storytelling toy from SuperModerno that never talks to the internet.

Show it a teddy bear and the toy turns that object into a bedtime story, narrated out loud. Its single big eye is the camera. A YOLOX-Nano detector looks at what you hold up, and only detections above 75% confidence count. A whitelist of supported objects keeps a stray mug in the background from becoming the hero of the plot.

How the story gets made

The detected object joins a profile a parent sets up in advance: the child’s name, hobbies, and interests. That prompt goes to a small language model running locally through llama.cpp. SuperModerno tried Gemma 3 1B and Qwen 3.5 0.8B, then kept Gemma 1B for the prototype. A local text-to-speech engine reads the result through the speaker behind the toy’s belly.

Two movable arms are wired to potentiometers, so a kid can bend the story by changing arm positions. The LED ring does double duty: it lights the object for the camera and plays an animation while the model thinks, so nobody wonders if the toy froze.

Where it went wrong

The gotcha list is useful reading. An untested USB power-and-data splitter cable fried the main board mid-build, which meant redoing part of the electronics. Speed was the bigger fight: early builds took up to 150 seconds before audio started. After refactoring the code and trimming the models, that dropped to about 27 seconds, still long for an impatient listener.

Try it yourself

  • Test every power cable with a multimeter before connecting your board. A bad splitter is a cheap lesson compared with a dead board.
  • Run a YOLOX-Nano or similar tiny detector on a laptop webcam first, and set the confidence cutoff to 0.75 to see how many false hits disappear.
  • Time your own pipeline stage by stage (detect, generate, speak) to find which one eats the seconds.

Read the original write-up on Hackster.io, then browse Arduino boards, cameras, and NeoPixel rings at circuit.rocks for a classroom edge-AI project.

Frequently Asked Questions

How does CantaStorie recognize objects without the cloud?

A YOLOX-Nano model runs on the Arduino UNO Q and only accepts detections above 75% confidence from a whitelist of supported objects.

What parts and skills does this build need?

An Arduino UNO Q, a camera, speaker, addressable LED ring, two potentiometers, and a battery pack. You also need basic Python, llama.cpp setup, and careful power wiring.

What will I learn if I build this?

You learn to run object detection and a small language model on-device, time and optimize an AI pipeline from 150 seconds down to 27, and wire LEDs and potentiometers as inputs. These skills fit edge-AI thesis and capstone projects.

This article was inspired by reporting from Hackster.io. Find the parts and modules to build it at Circuitrocks.

// written by Ann Arandia

Ann Arandia covers community projects and maker events for the Circuitrocks blog. She writes about local workshops, kid-friendly electronics, and the Philippine maker scene — the people, the meet-ups, the projects that come out of them.