AiMe is an
AI operating system.
One mind, 25 purpose-built specialists, scheduled and coordinated by a governing kernel — now running entirely on local hardware.
Most AI is one model behind a prompt. AiMe is an operating system.
The same problems an OS solves.
Take many independent components, give each a narrow job, schedule them, let them share memory safely, and compose their output into one trustworthy result — without any one of them going rogue. That is an operating system. AiMe is built like one.
25 specialists. Each narrow on purpose.
A fleet of small, mostly-local models and deterministic modules, composed into one coherent cognition. 24 of the 25 run live today — one (Expression) is intentionally held in reserve. Grouped by subsystem.
Green dot = live · amber dot = held in reserve. Three workhorses carry the fleet: one shared local gemma4:e4b serves eight roles at once, six custom-distilled models hold the sensing cluster, and a local gemma4:12b renders the final word.
Independence from rented intelligence.
The point of the fleet is that the cognitive stack is privately owned — not a thin client over someone else's API. As of June 2026, AiMe runs entirely on hardware its operator owns.
Custom in-house models
Small open bases — encoder heads and a few billion-parameter LLMs — distilled from multiple frontier teachers, then measured head-to-head against the very models they replace.
On shared local models
The rest run on shared local models on the operator's own hardware — fast, private, and free at the point of use. Eight roles share a single resident gemma4:e4b.
Remaining cloud call
Text-to-speech is the only outbound API on the live path — and it runs locally too, via Piper. Everything else: cloud models are strictly fallbacks, a safety net for a cold start, never the default path.
A continuous cognitive loop.
AiMe operates as a persistent loop — not a single request/response cycle. Each turn updates the user model and carries context forward.
Intent is classified deterministically before the LLM is invoked. The model is the last thing called, not the first.
Hippocampus RRF + Latent Episodes
calendar, scheduling, desktop
calendar, live feed
watch rules, pattern tracker
Live person-count, context snap
proactive candidates
It's a continuously maintained model of the user.
Context is not retrieved on demand — it's already active.
AiMe maintains a persistent, evolving model of the user — tracking identity, open concerns, relationships, behavioral patterns, and active goals. It holds these across sessions, not just within a single conversation.
The system doesn't look up facts about you when asked. It operates from a continuously maintained model of who you are — and everything relevant surfaces naturally from that context.
This is why swapping the underlying model doesn't break identity. The user model lives in AiMe — not in the model. Whichever cognitive engine responds, it responds from within the same persistent context.
Relevance by importance, not recency. Past context is scored by significance. High-significance episodes are injected before inference when the current turn connects to established portrait content.
A persistent model of the user. Six layers deep.
AiMe builds and maintains a structured, evolving representation of the user — persisted across sessions, updated after every turn. What they care about. What concerns remain open. What patterns repeat. What context is currently active.
Direct routing. Every specialist.
Intent is classified pre-LM by a governed specialist. The v4 spine runs sequential specialist calls — each contract-bound to a single role, each with its own model selection. The whole spine now runs local-primary, with cloud kept only as a resilience fallback.
The spine runs local-primary end to end — the final renderer is a local gemma4:12b — with cloud providers (Ollama · Gradient · xAI · Gemini) kept only as a resilience fallback. No provider outage silences the system, because the default path never leaves the machine.
Six sovereign agents.
Six independent agents — each with its own state, lifecycle, and operational scope. They surface to AiMe. They do not narrate. They never block the turn path.
ProactiveLoop with ambient triggers and DMN scheduling. Multi-turn event staging. Conflict detection. Proactive schedule-candidate pipeline from email and chat. Explicit user directives auto-commit — suggestion paths stay approval-gated. Return and morning brief generation.
Significance-filtered inbox management. Multi-provider — Gmail OAuth and IMAP/SMTP. High-significance unread emails surface pre-LM as ★-marked entries. Behavioral feedback loop adjusts scores.
Vision-guided desktop automation. Plan → Confirm → Execute protocol. Live presence via Haar cascade (~200ms). Face-count delta triggers arrival and departure turns. 120s exit grace compensation. Fallback to snapshot DB when Thalamus unavailable.
Sovereign file agent with Windows-aware sandbox covering 11 threat vectors. Fingerprinted plan → confirm → execute with confirmation protocol. Atomic writes with rollback. Content-smuggling defense. 88 tests, 8 review rounds.
Multi-executor coding orchestration. Coordinates code writer, test runner, linter, and file manager executors. Structured task decomposition with review discipline built into the execution loop.
All standing interests — imprints, watch rules, and portrait concerns — are indexed as 768-dim BGE vectors in Qdrant. Every observation is checked against this index. Matches surface through a 6-check companion filter (quiet hours → activity gate → rate limit → per-intent cooldown → semantic dedup → consent grade) before AiMe is interrupted. Pattern tracker escalates recurring signals through 4 levels with significance boosts.
AiMe initiates turns — not just responds. 5-tier absence grading determines re-engagement tone. Return recognition fires a local-vision snapshot on arrival. Third-party presence detected via live bbox-count delta: count increase triggers an arrival turn, count decrease (while others present) triggers a departure acknowledgment. Fresh-boot detection prevents spurious greetings on system restart.
Same character. Different frame.
A swappable identity layer prepended to an immutable core operational contract. Switching personas changes voice and relationship stance — not the truth architecture or memory system beneath.
Personal companion mode. Relationship-first, continuity-aware. The Living Portrait of the user is the primary context frame. Warmth within the operational contract.
Operational governor mode. System Portrait active. Role-keyed authority bounds, incident stack, governance commitments. The model responds to role, not identity.
Simultaneous home and work frames. Same cognitive substrate, two portrait subjects. Context-aware blend: one frame when only one is active, both when both are relevant.
Production components
Live operational status of AiMe's core subsystems as of June 2026 — running entirely on local hardware.
AiMe does not wait for prompts.
AiMe monitors relevance over time and surfaces information when it becomes meaningful. A concern mentioned days ago can reappear when conditions change. A pattern can be recognized without being explicitly asked. This is behavior driven by continuity — not input alone.
A system that remembers — without being asked.
Because it was still relevant.
Models are interchangeable. Continuity is not.
AiMe sits above any underlying model. What the model provides is inference. Everything else — identity, memory, behavior — lives in AiMe.
Continuity is preserved across model updates, provider switches, and capability upgrades — by design.
Request access to AiMe.
AiMe is currently in private deployment. Tell us about your use case and we'll follow up directly.
The orchestration is.