Project AURA
A delusion made real.
A real-time AI companion with a sub-second voice loop, an emotion-aware Live2D rig, and a control center that hot-swaps the entire AI brain at runtime.





Five distinct emotion states drive a Live2D rig in real time.
Every LLM reply gets classified into one of five emotion states: happy, playful, surprised, companion, or idle. The avatar's Live2D parameters (eyes, mouth, head tilt) blend toward that state while the TTS audio plays. Companion mode also summons the ghost helper.
detected from positive sentiment + exclamation tokens
The Pipeline
four hops, under 800msDemo
v0.1.1 · live voice loopThe Emotion System
five poses, blended over Live2D parametersAvatar Rendering
browser-native, zero external softwareThe avatar runs entirely in the browser, with no VTube Studio and no external process. The animation loop is monkey-patched into coreModel.update() so every frame gets injected parameters right before GPU commit.
Organic Blink
A finite state machine fires random timers for the blinks. Probability and cooldown both vary, so the pattern never repeats exactly.
Eye Saccades
Every frame adds a little randomization to gaze direction via Math.random(), so the eyes drift the way real eyes do and never sit perfectly still.
RMS Lip Sync
AnalyserNode computes Root Mean Square amplitude from the LiveKit audio track in a requestAnimationFrame loop and maps it directly to mouth open.
AI Memory & RAG
reads documents, recalls on demandTeaching AURA
Upload any document. AURA chunks and embeds it, and the contents are searchable in the middle of a live conversation.
AURA Remembers
Every utterance triggers a semantic search. Matching chunks get injected into the LLM context with no visible step, so AURA just seems to know.
The Interface
the avatar plus the whole control surfaceWhat I built · what's next
What's in v0.1.1
- ✦End-to-end voice loop under 800ms from speech end to TTS start.
- ✦Multilingual in and out. Deepgram and Qwen 3 handle EN/ID switching mid-sentence.
- ✦Five-state emotion system with prompt-based classification per turn.
- ✦Live2D rig driven by emotion vectors, lip-sync from TTS phonemes.
- ✦Hot-reload model swap to change LLM provider or model live from the control center.
- ✦Per-context memory so every chat keeps its own personality, history, and creativity dial.
What's next
- →Tool use via function-calling for calendar, search, and code execution.
- →Long-term memory with vector recall across contexts.
- →Vision input from webcam frames, so AURA can see what you're working on.
- →Avatar marketplace with pluggable Live2D rigs and rig-aware emotion mapping.
- →Mobile client for the conversation loop.
- →v1.0 when emotion classification moves on-device.
Building real-time AI? Let's compare notes.
I'm always up for swapping notes on Live2D rigging, sub-second voice pipelines, or emotion classification prompts. And if you want AURA's roadmap to head somewhere specific, tell me.