Empirical Hardware Benchmarks
Real forensic measurements executed against physical hardware (17.18 GB Unified Memory Host) across 9 simultaneous multimodal pressure scenarios and micro-benchmark test harnesses.
The 9 Multimodal Pressure Scenarios
Select a scenario to inspect measured RAM consumption and active concurrency contention.
System Idle
Full container mesh running (Postgres, Neo4j, Redis, Qdrant, NATS, LiveKit) with no active conversation.
Subsystem Micro-Benchmarks
Exact timings captured from test suites and live measurement runs in tools/measure/out/.
Composed Turn Time-To-First-Token (TTFT)
Composed turn deliberation + prompt prefill on dedicated RTX 2060 Super (isolated prompt eval: 39.95 ms)
RTX 2060 Super (3B) Generation Speed
Sustained streaming throughput (187ms full sentence generation, zero audio buffer underrun)
Hermes 3 (8B) Time-To-First-Token
Empirical streaming generation benchmark across 5 companion scenarios on Tesla T4 GPU
Hermes 3 (8B) Generation Speed
Sustained streaming throughput (~150ms per 7-word audio chunk, zero TTS buffer underrun)
Sub-Millisecond Barge-In Reflex
Audio interruption reflex dispatch latency over NATS bus upon voice activity detection
Deterministic State Rollback
Atomic rollback of speculative cognitive state mutations upon turn cancellation
Identity & Safety Boundary Invariance
Constitutional guard evaluation across 1,200 adversarial redteam probes
Metacognitive Latency Overhead
Calibration and second-order reflection directive generation overhead
Candidate Action Selection Latency
Constraint-first scoring and selection across 10 competing action intents
Deterministic Plan Verification
Formal DAG verification of multi-step deliberative plans prior to execution
Neo4j Knowledge Graph Warm Fetch
1,003 entities & 4,002 relations fetched with 300s TTL memory cache
Neo4j Knowledge Graph Cold Query
Full Cypher graph traversal over 1,003 un-cached entity nodes
Neo4j Graph High-Volume Seeding
Concurrent ingestion of 1,000 entity nodes and 1,998 relationship edges into graph DB
Subconscious REM Consolidation (Idle)
Extracts facts and updates Neo4j beliefs across 6 recent turns
Subconscious REM Consolidation (VLM Load)
Consolidation pass executed during continuous Moondream VLM inference
LiveKit Audio Frame Burst Delivery
50 sequential 32kHz PCM audio frames published over NATS
Docker Mesh Memory Breakdown (752.4 MiB Total)
Measured resident container memory across the 6 core backend services.
Platform Performance Matrix
Engineering targets the architecture is designed for, not live measurements — see Section 2 above for what's actually been measured.
| Hardware Target | Launch Profile | LLM Inference Engine | Voice Engine | Time-To-First-Token | Total Turnaround |
|---|---|---|---|---|---|
| Dedicated Home GPU (NVIDIA RTX 2060 Super 8GB) | Headless 24/7 Full Mesh (Linux Kernel 7.0 / Ubuntu 24.04) | Llama 3.2 3B (Ollama Bare-Metal / CUDA 13.2) | GPT-SoVITS 32kHz (CUDA) | 119.35 ms composed (39.95 ms isolated prompt eval) | 120 - 180 ms |
| Google Colab / Cloud GPU (NVIDIA Tesla T4) | Full mesh + Hermes 3 (8B) | Hermes 3 8B (Ollama / CUDA) | GPT-SoVITS 32kHz (CUDA) | 61.9 ms (Measured) | 160 - 220 ms |
| Apple Silicon M1 / M2 / M3 (16GB Unified) | Full Stack (Voice + Brain + Memory) | Llama 3.2 3B (Metal / MLX) | GPT-SoVITS (CPU / Metal) | 320 - 450 ms | 680 - 950 ms |
| NVIDIA RTX 3060 / 4060 (12GB VRAM + 16GB Host) | Full Stack + Vision Profile | Hermes 3 8B / Qwen 2.5 14B (CUDA) | GPT-SoVITS 32kHz (CUDA) | 50 - 90 ms | 150 - 260 ms |
| Modern x86_64 CPU (16GB RAM, No GPU) | Heavy Mode (Local Whisper STT + Brain) | Llama 3.2 1B (AVX-512 / OpenVINO) | Bundled Pre-synthesized Reference | 450 - 650 ms | 900 - 1400 ms |
| Weak Laptop + Cloud Fallback (8GB RAM) | Light Mode (Claude 3.5 Sonnet Fallback) | Anthropic Claude API (Streaming) | Remote TTS or WebRTC Voice | 280 - 400 ms | 600 - 850 ms |
Conversational Loop Latency Waterfall
A per-stage budget, not a measured trace — the direct attempt to measure this loop (m14_stt_cost.json) came back UNKNOWN because stt-agent wasn't running for that run.
Speech Detection & VAD Cutoff
Energy-based thresholding & silero VAD
Speculative Intent & Emotion
Early barge-in reflex & 7-class emotion classification
Final Speech Transcription
High-precision word-level transcript generation
Appraisal & Endocrine State
PAD computation, boundary check, cortisol/dopamine update
Deliberation & Intent MAUT
Behavior tree traversal and candidate scoring
LLM Time-To-First-Token (TTFT)
Empirical streaming first token dispatch (61.9ms measured on Tesla T4)
GPT-SoVITS 32kHz Synthesis
Streaming chunk synthesis with prosody trajectory & pause bias
LiveKit WebRTC Transmission
PCM audio frames + visemes data channel dispatch