A friend
of your own
making, on your
own machine.

Describe them in your own words. They remember who you are through ACT-R-scored recall, build trust and attachment through real appraisal math, and reason through a 7-stage cognitive loop before answering in a cloned voice you gave them — running entirely on your own hardware, 100% local, MIT licensed.

100% Local-FirstMIT LicensedACT-R Memory DecayTonic/Phasic EndocrineNATS JetStream Mesh
NEUROCHEMICAL SIMULATION

Mood that changes how
the LLM generates words.

Cortisol and Dopamine aren't decorative strings — they mathematically modulate LLM sampling temperature, top-p, and token ceilings in real time.

Affective Modulation & Sampling Physics

Endocrine Parameter Modulator

Adjust simulated neurochemicals to see how internal affective states dynamically reshape LLM generation.

Cortisol (Stress / Vigilance)0.30
Calm / RelaxedHalf-life: 4500sHigh Alert / Defensive
Dopamine (Reward / Engagement)0.50
NeutralHalf-life: 90sHigh Enthusiasm
Physical Fatigue (Turn Fatigue)0.20
EnergizedAccumulates over turnsExhausted
Inferred LLM Sampling & Acoustic Parameters
LLM Temperature0.72

Creative & Broad

Sampling Top-P0.82

Tight Lexical Selection

Max Response Tokens220

Expansive Explanations

Estimated Tempo159 WPM

Physical speech pacing & pause bias

Affective Independence: Phasic bursts decay according to calibrated half-lives, allowing the agent to be stressed by an urgent task while enthusiastic about solving it.
THE COGNITIVE ENGINE

Sub-second reflex &
deliberative appraisal.

Step through how raw 16kHz audio moves through the 9-agent asynchronous NATS signal bus from speculative emotion classification to 32kHz neural voice rendering.

Cognitive Pipeline

The 7-Stage Cognitive Turn

Step through how a single human utterance traverses the decentralized multi-agent mesh.

01 / 07

Perception

Ingests raw 16kHz PCM audio stream from user microphone and publishes to NATS audio.inbound.

Publisher Agenttransport_agent
Engine / TechnologyLiveKit WebRTC
NATS Subjectaudio.inbound
Target Latency15 msDesign budget, not a live measurement
FREEFORM PERSONA COMPILER

Describe them in prose.
Never pick from a preset list.

Our compiler translates natural language into a strictly enforced 3-tier constitution (Immutable Safety Core, Constitutional Temperament, Adaptive Traits) with explainable parameter mappings.

Live Simulator

Persona Compiler: The Real Math

In the full system, an LLM reads your prose and scores it on 9 dimensions. Here you set those 9 scores directly with the sliders below and see the exact deterministic formula (_infer_temperament) that turns them into a persona — computed live in your browser, not simulated.

Warmth0.70
ColdWarm
Energy0.57
CalmExcitable
Assertiveness0.64
YieldingTake-charge
Volatility0.35
Even-keeledReactive
Resilience0.90
Dwells on thingsBounces back
Opinion Firmness0.85
Easily swayedStubborn
Openness to Trust0.85
GuardedQuick to trust
Warmth Growth0.70
StandoffishQuickly attached
Emotional Lingering0.50
Brief reactionsLingering reactions
Baseline Valence0.42
warmth=0.70 -> how warm vs. cold the description reads, scaled into the ±0.6 valence bound (a friend can never be pinned fully positive)
Baseline Arousal0.549
energy=0.57 -> calm/low-key (0) to excitable/high-energy (1), scaled into the 0.15-0.85 bound
Baseline Dominance0.598
assertiveness=0.64 -> yielding (0) to take-charge (1), scaled into the 0.15-0.85 bound
Valence Drift Rate0.31
volatility=0.35 -> how much mood swings drives how fast valence itself moves
Arousal Response Rate0.377
volatility=0.35 -> a more reactive temperament also means arousal responds to events faster
Dominance Stability0.645
opinion_firmness=0.85 -> easily swayed (0) to stubborn/consistent (1)
Trust Change Rate0.39
openness_to_trust=0.85 -> guarded (0) to quick-to-trust (1) shapes how fast trust itself can move
Attachment Growth Rate0.195
warmth_growth=0.70 -> standoffish long-term (0) to quickly-attached (1)
Mood Decay Rate0.38
resilience=0.90 -> dwells on things (0) to bounces back quickly (1) -- higher means faster return to baseline mood
Dopamine Half-Life (s)180
emotional_lingering=0.50 -> how long a good moment's glow lasts, in seconds
Cortisol Half-Life (s)700
emotional_lingering=0.50 -> how long a bad moment's sting lingers, in seconds -- longer than dopamine's by construction
Adrenaline Half-Life (s)350
emotional_lingering=0.50 -> how long a startle/interruption/shock reaction lingers -- sits between dopamine's and cortisol's by construction
Initial Trust0.71
openness_to_trust=0.85 -> where the relationship's trust starts (never at the extremes)
Initial Attachment0.295
warmth_growth=0.70 -> where attachment starts -- deliberately low; attachment is meant to be earned
These 14 values are computed fresh on every slider move by the exact same formula as backend/app/persona/compiler.py. Try the Trust & Attachment demo to see how trustChangeRate and attachmentGrowthRate above then drive a relationship forward, turn by turn.
MARSH TRUST + BOWLBY ATTACHMENT

Trust that's earned,
not scripted.

Every appraisal updates benevolence, competence, and integrity independently, and attachment grows on a slower, frequency-gated curve than trust does.

Live Simulator

Trust & Attachment Over Time

Exact port of AgentState.update_from_appraisal — Marsh (1994) trust (benevolence/competence/integrity) and Bowlby attachment, both driven by the same appraisal that updates mood. Set the appraisal for one simulated turn, advance it, and watch the relationship move.

Goal Congruence0.50
Did this help what you're trying to do?
Relationship Impact0.30
Did this feel like a moment between you two?
Novelty0.20
How unexpected was it?
Relevance0.50
How much did it matter right now?
Agency0.40
How much control did you have over it?
Norm Alignment0.60
Did it match what's expected/appropriate?
0 turns simulated
Benevolence0.500
Competence0.500
Integrity0.500

Attachment grows slower than trust by construction: it's scaled by min(1, interaction_count/100), so even a friend who trusts you immediately still needs 100 more simulated turns before that frequency term stops suppressing attachment growth.

ACT-R BASE-LEVEL ACTIVATION

Memory that decays
like the real thing.

Recency, frequency, importance, and emotional proximity combine into a single recall score — the exact formula every retrieval path in the backend shares.

Live Simulator

Memory Activation & Decay

Exact port of the ACT-R base-level activation formula shared by every retrieval path in the real system: ln(freq) − d·ln(recency) + importance + emotional-proximity + spacing. The graph below is illustrative fixture data (not a real conversation), but the activation score on every node is computed live from this formula.

Recall Count3
Hours Since Last Recall48h
Importance Score0.60
Emotional Distance (0=matches current mood)0.40
Massed vs. Spaced Recall

Same frequency and recency — spaced practice still wins, the literature's central spacing-effect finding.

Illustrative Memory Graph (click a node to query it)
workfamilyhobbiesfriendshipself

Brightness is each node's live ACT-R activation. Clicking a node simulates querying it: its direct neighbors (amber ring) get a one-hop spreading-activation boost — a simplified stand-in for the real system's Personalized-PageRank graph boost, not a literal port of that iterative algorithm.

HOW THE BRAIN'S AGENTS COORDINATE

Separate processes,
not function calls.

Agents coordinate over NATS JetStream with typed Pydantic contracts — a real signal-bus mesh, all running locally.

BRAIN AGENT
BRAIN AGENT

The cognitive core

Runs the full turn: appraises what happened, updates PAD affect and the endocrine layer, decides how to respond, and streams the reply through your local Ollama model.

Python
runtime
Ollama
LLM backend
VOICE AGENT
VOICE AGENT

Speech, physically synthesized

Renders every reply through one self-hosted, cloned voice — no fallback to a different voice. Pauses, ducking, and prosody are real PCM sample manipulation, not text markers.

Rust
runtime
GPT-SoVITS
synthesis
STT AGENT
STT AGENT

Dual-path listening

A fast speculative path classifies intent and emotion mid-sentence for natural barge-in; a slower accurate path produces the final transcript once you're done talking.

whisper.cpp
accurate path
SenseVoice
fast path
SUBCONSCIOUS AGENT
SUBCONSCIOUS AGENT

Reflection, between turns

Consolidates memory, decays what's no longer relevant, and can reach out to you unprompted — a friend that thinks about you when you're not talking to it.

Neo4j
knowledge graph
ACT-R
memory decay
Get started

Four commands.
That's the whole flow.

terminal
# Clone and configure
$ git clone https://github.com/PALabs-v1/AI_friend.git
$ cd AI_friend && cp .env.example .env
# Network, Ollama, default voice, schema, mesh — all of it
$ ./start.sh
==> Starting the mesh...
==> Done.
VOICE — A SUPPORTING MODALITY

Studio-quality 32kHz voice,
cloned from 8 seconds.

The Brain drives what gets said and how it's said; GPT-SoVITS renders it with dynamic pause-bias scaling and sub-millisecond barge-in (<1ms).

Reference Parameters — No Audio Yet

Emotional Voice Cloning & Prosody

The real per-affect reference-clip and pause-scale mapping the voice agent uses. No rendered audio is wired up on the website yet — see the roadmap.

Sample line for this register
"Take your time. There's no rush to figure this out tonight."
Affect Parameter SpaceLow Arousal, Positive Valence
Pause Scaling Bias1.25x (Relaxed cadence)
Acoustic Characteristic

Generous inter-phrase pauses and lowered pitch variance, suitable for deep reflection and late-night conversation.

Engine: GPT-SoVITS (32,000 Hz)Barge-in reflex: < 1ms (0.099 ms validated)
EMPIRICAL MEASUREMENTS

Measured on physical hardware.
No ungrounded claims.

Detailed memory footprints, latency breakdowns, and peak load stress test verification.

Empirical Hardware Measurements

Hardware Matrix & Latency Waterfalls

Every figure measured on live physical hardware — no theoretical estimates presented as results.

Explore All 9 Scenarios →
Hardware PlatformOperational ProfileLLM Time-To-First-TokenTotal Loop TurnaroundStatus
Dedicated Home GPU (NVIDIA RTX 2060 Super 8GB)Headless 24/7 Full Mesh (Linux Kernel 7.0 / Ubuntu 24.04)119.35 ms composed (39.95 ms isolated prompt eval)120 - 180 msVerified On-Premises Hardware
Google Colab / Cloud GPU (NVIDIA Tesla T4)Full mesh + Hermes 3 (8B)61.9 ms (Measured)160 - 220 msVerified Cloud Telemetry
Apple Silicon M1 / M2 / M3 (16GB Unified)Full Stack (Voice + Brain + Memory)320 - 450 ms680 - 950 msSupported Local Baseline
NVIDIA RTX 3060 / 4060 (12GB VRAM + 16GB Host)Full Stack + Vision Profile50 - 90 ms150 - 260 msUltra Low Latency Tier
Modern x86_64 CPU (16GB RAM, No GPU)Heavy Mode (Local Whisper STT + Brain)450 - 650 ms900 - 1400 msSupported Baseline
Weak Laptop + Cloud Fallback (8GB RAM)Light Mode (Claude 3.5 Sonnet Fallback)280 - 400 ms600 - 850 msCloud Hybrid Tier
Turnaround Latency Composition (Speech → Cognitive Appraisal → 32kHz Audio Output)
15 msSpeech Detection & VAD Cutoffstt-agent (Rust)
135 msSpeculative Intent & EmotionSenseVoice (sherpa-onnx)
180 msFinal Speech Transcriptionwhisper.cpp (Rust)
4 msAppraisal & Endocrine Statebrain_agent (Python)
SECURITY & PRIVACY

Local by design,
not by promise.

No cloud accounts, zero telemetry, and atomic 4-store disaster recovery portability.

Data Sovereignty

Security, Privacy & Local Confinement

Designed for developers and researchers who demand verified local containment without cloud data leaks.

ZERO EGRESS

100% Local Confinement

PostgreSQL, Neo4j, Redis, Qdrant, NATS, Ollama, and LiveKit bind exclusively to 127.0.0.1 loopback interfaces.

Zero conversation text, audio, or vector embeddings leave your hardware.
LEAST PRIVILEGE

RBAC JetStream Isolation

Every NATS client operates under scoped user accounts with restricted publish/subscribe topic permissions.

Prevents unauthorized agents from reading sensitive system subjects.
PORTABILITY

4-Store Atomic Snapshots

Single-command export script captures PostgreSQL JSONL, Neo4j Cypher, and SQLite affect state into an encrypted archive.

Total disaster recovery and seamless migration across workstations.
CONSTITUTIONAL

Immutable Safety Floor

Tier 0 hardcoded boundaries (Honesty, Privacy, Anti-Harm) enforced before any user prompt can reach the LLM.

Guarantees system credentials and private files cannot be exfiltrated.
NO TELEMETRY

Zero Account Tracking

The entire codebase contains zero tracking scripts, analytics pixels, or phone-home logging.

Complete anonymity with true self-hosted ownership.
OPEN SOURCE

Permissive MIT License

Full transparency under the MIT license with public audit reports, mutation tests, and CI workflows.

Freedom to inspect, fork, modify, and extend without commercial lock-in.

No waitlist.
It's already yours to run.

Free and open source, MIT licensed. Clone it, describe your friend, and start talking — no account required.

curl -fsSL https://raw.githubusercontent.com/PALabs-v1/AI_friend/main/scripts/install.sh | bashREAD INSTALL GUIDE