VIRTUAL TWILIGHT
ENGINEERING PROOF · MEASURED JUN 2026
◢ BUILD INTEGRITY: VERIFIED
FROM CHATBOTS TO MINDS

You can't vibecode a living world.

Virtual Twilight isn't a persona card bolted onto an LLM. It's 1.09 million lines of production code, a 66-step AI pipeline with 7 parallel execution groups, 35+ modular prompt sections assembled per message, 5-provider LLM routing with a background-cost engine, and nine psychological simulation systems that permanently change with every conversation. The 1,031 throwaway debug scripts are the fossil record of the thinking that built it.

1.09M
Production lines
1,296
API endpoints
66
Pipeline steps
35+
Prompt sections
5
LLM providers
01 /

Lines of code — real, measured

Counted directly from the repository, June 2026. Excludes dependencies, build output, and caches. Most vibecoded "AI apps" are 2,000–10,000 lines. The frontend alone is 59× that.

Part of the systemFilesLines
Frontend web/src · TS/React1,304~590,000
Backend services app/services588~357,000
Backend API layer app/routers · 1,296 endpoints156~115,000
Database layer app/database · 30+ models62~27,000
Production core subtotal~2,100~1.09 million
Full backend total app/4,001~1.22 million
Audit / debug / investigation scripts _*.py — fossil record of thinking1,031~132,000
02 /

The 66-step pipeline — every single message

66 discrete named steps counted directly from facade.py, June 2026. 47 execute before or immediately around the AI model call. 19 fire after the reply is sent — permanently updating personality drift, memory, relationships, and achievements. Not marketing; it's a counted list that has grown as the system has grown.

Group A — Security

  • input_security — injection / jailbreak detection
  • knowledge_boundary — hard-limit policy enforcement

Group A1 — Data Loaders (parallel)

  • load_npc_structural_relationships
  • assistant_features — web search, remember, recall (Brave API + pgvector)
  • load_relationship_data — 12-dimension trust matrix
  • load_conversation_history
  • load_physical_attributes

Group A2 — Mid-pipeline (parallel)

  • proactive_communication — NPC-initiated outreach
  • load_memory — pgvector importance-scored recall
  • knowledge_step — semantic vector search
  • unified_player_analysis — VAD sentiments, intentions, 240 behavioral templates

Group B — World State (parallel)

  • location_data — ELS sub-location adjacency graph
  • collective_knowledge — world lore, events
  • item_context
  • player_fact_extraction
  • player_cross_npc_awareness — gossip network
  • training_scenario
  • coach_knowledge / coach_conversation_context

Group C–D — Rules, Emotion & Intent (parallel)

  • project_world_rules
  • time_awareness
  • personality_seeding — Big Five shaping
  • npc_emotion_estimate — VAD 3-axis physics
  • calculate_relationship_metrics
  • proximity — spatial distance detection
  • player_promise_capture — detect "I'll / I promise" for callback
  • promise_lifecycle — resolve kept/broken/overdue promises
  • communication_style — response length/style guidance
  • clarity_check — non-native-speaker NPC clarification gate
  • trigger_awareness — pre-response keyword behavioral scan
  • combat_intent — hostile NPC context injection (gated)
  • internal_reflection — silent subtext layer before prompt build

Prompt Build → LLM → Output

  • enhanced_generate_prompt — 35+ modular sections, priority-ordered, shrinkable
  • lora_activation — character LoRA selection + injection
  • response_executor — Azure GPT-4o / Venice / Nano / Gemma (280–450ms TTFT)
  • training_goal_validation — rollback goal completion on AI failure
  • intimate_reaction — inject physical reactions if NPC omitted them
  • output_sanitizer — 3-layer prompt-leak defense
  • training_authenticity_regen — detect bland responses in uncensored mode, regenerate
  • npc_action_generator — optional physical action separate from dialogue
  • health_drain — species-specific drain during fight/intimate
  • response_validation — Voight-Kampff: no "I am an AI" leakage
  • npc_item_give — detect item-giving in response, create+transfer to player
  • post_response_self_critique — NPC flags self-critique signals for next turn
  • finalize_response — format final payload

19 Post-processing Steps — fire after reply is sent

  • location_change — handle movement triggered by response
  • appointment_extraction — detect date/meeting agreements, create records
  • process_emotion — consolidate emotion state pre-TTS
  • speech_integration — TTS generation after emotion processing
  • persist_emotion — VAD → npc_emotion_state DB
  • process_relationship — trust/familiarity drift
  • achievement_tracking — content counters
  • gamification — check achievement thresholds
  • emotion_shock — permanent personality drift
  • conversation_persist — save messages
  • update_interaction_count
  • event_step — trigger evaluation
  • memory_step — pgvector importance-scored write
  • motivation_memory — drive setpoint update
  • milestone_anchor — promote first_meeting / first_gift memories to protected anchors
  • metrics_step
  • emotion_broadcast — WebSocket push to Unity/React
  • post_session_analytics — self-learning data write
  • finalize_response — format final payload (no sanitization)
03 /

The modular prompt builder — 35+ sections per message

The LLM never sees a hand-written prompt. Every message triggers a context-aware assembly engine that selects, prioritises, shrinks, and orders up to 35 named prompt sections — then routes the output to the cheapest model that can handle it.

🔒

CRITICAL_ALWAYS_ON (~12 sections)

Never filtered regardless of intent or mode: anti_hallucination, output_contract, response_behavior, core_identity, character_behavior_rules, behavior_instructions, user_priority, character_profile, temporal_anchor, triggered_events. These form the cache anchor target.

🧱

CORE (~25 sections)

Always included regardless of intent confidence (changed Jun 2026 — no more minimal_core): unified_psychology, physical_attributes, conversation_history, enhanced_location, time_context, memory_facts, internal_thought, self_reflection, motivation, player_model, self_direction + more.

⚡

RULES[intent] + mode overrides

Intent-keyed sets add flavour on top. ACTION MODE drops RPG/companion sections and boosts action-relevant ones. TRAINING MODE drops RPG/companion. Knowledge section conditional on knowledge_count>0. Section set changes per-turn.

Prompt cost breakdown — measured from production (NPC 2718, 35 blocks, ~16K tokens)

knowledge block · ~2,600 tok · BIGGEST · every turn conversation_history · ~2,200 tok psychology · ~1,550 tok world_locations · ~1,490 tok final_behavior · ~1,274 tok emotion_state · only ~70 chars · likely under-invested

Cache potential: gpt-4.1-mini gives a 50% discount on cached input tokens (≥1024 identical prefix). Currently ~0% effective because two "static" sections have per-turn content gates that flicker them in/out, breaking the byte-identical prefix. When fixed: a CACHE ANCHOR of always-on, content-invariant sections pinned to the front could halve input cost on ~96.7% of all dialogue traffic.

04 /

The LLM is step 38 of 66 — the backend does the thinking

Every competing AI agent platform treats the LLM as the intelligence. ELMU treats it as the voice. 37 steps run before the LLM call, computing the full psychological state, memory recall, world context, and knowledge pack retrieval into a ~16K-token context. The LLM generates ~200 tokens. Then 28 post-process steps update personality, memory, emotion, and relationships. The LLM can be swapped (Azure → Mistral) without changing agent behaviour — because the intelligence is not in the LLM.

⚙

Steps 1–37: The backend computes

Before the LLM ever speaks: security + boundary → NPC identity + cache → relationships → conversation history → pgvector memory recall → player VAD analysis → promise lifecycle → world state + lore → knowledge pack retrieval → Big Five personality seeding → VAD emotion estimate → relationship metrics → 35-section prompt assembly → LoRA activation.

The LLM receives a fully computed ~16K-token context. It does not decide personality, emotion, memory, or relationship state. The backend already did.

🗣️

Step 38: The LLM speaks

ResponseExecutorStep — one step out of 66. Currently routes to cheapest model that can handle the computed context:

NOW (5 providers):

  • Azure GPT-4.1 — primary dialogue
  • Azure GPT-4.1-nano — fast/cheap SFW narration
  • Venice/Dolphin — uncensored, TWILIGHT mode
  • Gemma-4-uncensored — background work (256K ctx)
  • Mistral+LoRA — dev / fine-tuned characters (CINECA)

TARGET (3-model Mistral stack, EU GPU):

  • Mistral Large fine-tune — primary dialogue, all projects
  • Mistral 7B instruct — narration + all background tasks
  • Mistral 7B uncensored LoRA — TWILIGHT / mature mode
🔄

Steps 39–66: The backend updates

After the LLM speaks: output sanitisation → authenticity check → Voight-Kampff validation → emotion processing → TTS → persist VAD → relationship drift → achievement tracking → emotion shock (permanent personality drift) → pgvector memory write → motivation memory write → milestone anchor protection → WebSocket emotion broadcast → self-learning data write.

The NPC at message 600 is measurably different from message 1 — not because the LLM changed, but because 28 steps permanently updated the state the LLM will read next time.

The EU responsible AI training flywheel

Every interaction produces a (computed_context, response) pair. Context = 37 steps of psychological state. Response = what a psychologically-grounded agent said given that state. These pairs, anonymised via EAL before leaving the inference path, are the fine-tuning dataset stored on CINECA Leonardo (EU HPC, Italy). Fine-tuned Mistral learns to speak consistently with the psychological models — because it was trained on outputs generated after running them. The LLM slot becomes progressively more accurate, cheaper, and faster, without ever changing the pipeline that supervises it.

better pipeline → better training pairs better Mistral → better responses more tenants → more training data EAL anonymisation → GDPR-safe at capture CINECA storage → EU data sovereignty OVH/Scaleway hosting → no US dependency in loop
05 /

Nine psychological simulation systems

These are not features on a feature list. Each is a separate engine with its own data model, its own persistence layer, and its own permanent effect on NPC state. They interact with each other — which is exactly the kind of thing you can't generate without designing.

⬡

VAD Emotion Engine

Three continuous axes — valence ↔, arousal ↕, dominance ↕ — updated every message. Emotions decay over time like real feelings. Emotional shocks (betrayal, intense intimacy, humiliation) cause permanent personality drift with irrationality modifiers: amplified, dampened, inverted, or longing patterns.

◎

Big Five Personality

Openness · Conscientiousness · Extraversion · Agreeableness · Neuroticism — each scored 0–100. Shapes every downstream behaviour: how fast they trust, what makes them uncomfortable, fighting style, memory what they fixate on, how emotions escalate. Not cosmetic — these are computation inputs.

◉

9-Drive Motivation System

Safety · Social · Competence · Status · Curiosity · Autonomy · Meaning · Intimacy — each with baseline, importance weight, and decay rate. Drives generate autonomous behaviour when unmet. When chronically unmet, they permanently raise the NPC's setpoint. When satisfied repeatedly, they lower it.

◇

Infinite Persistent Memory

Not a context window — a pgvector database. Importance-scored, personality-filtered recall. A jealous NPC remembers every mention of a rival. A forgiving one lets slights fade. Memory grows indefinitely with no reset. Scales horizontally with the database — no architectural ceiling.

◈

12-Dimension Relationship Matrix

Trust · Chemistry · Intimacy · Rapport · Compatibility · Stability · Familiarity — all tracked independently across every NPC–player pair. Trust earned over weeks drops in one sentence. Betray one NPC and watch it ripple through the gossip network to others.

⬢

Autonomous Life System

NPCs move between sub-locations via an adjacency graph (Enhanced Location System). They gossip about each other, spread rumours, send messages when you're gone, and generate daily chronicles. This runs whether you're online or not. The world doesn't wait for your message.

🌙

Sleep · Dream · Psyche Loop

NPCs sleep on a schedule. During sleep, 9 dream types fire: creative, shared, nightmare, anxiety, cascade, prophetic, memory_replay, connection, desire. Dreams write to memory, promote insights to the knowledge base, and cause permanent psyche drift — the NPC wakes slightly different. 19,492 dream actions verified in production DB.

🧭

NPC Self-Agenda & Agency

Once autonomy + competence drives accumulate enough positive drift, NPCs form a self-directed 4-step goal plan. Self-sufficiency score rises with pursuit count (0.0→1.0). At 0.90: "I don't need them the way I once thought. I've figured out something better — my own path." At 0.75: they physically walk out of the location on their own (not from hostility — from having a life).

🛠

Embodied Skill Learning

NPCs accumulate craft skills (dexterity from practice, movement for world-aware NPCs) with diminishing returns: gains shrink near mastery, losses cost more at high levels. Motor realizations surface in the prompt as body-felt first-person thoughts. The same loop is the literal hook for physical robot motor learning — nothing in the chain changes for a real body.

06 /

Why "infinite memory" is an engineering claim, not a marketing claim

✗ Context-window memory

  • Fixed token ceiling (16K, 128K — doesn't matter, it's a ceiling)
  • Resets or summarises per session
  • No personality filtering — everything is equally memorable
  • No semantic search — recency wins, not relevance
  • Memory can't grow without re-engineering the architecture

● Database memory — no architectural ceiling

  • pgvector with semantic similarity search — recall what matters, not what's recent
  • Importance-scored at write time — personality shapes what sticks
  • Grows indefinitely. Scales with horizontal database replicas.
  • Cross-NPC: one NPC can recall something another told them
  • Dream-promoted insights elevate memory to permanent knowledge base entries
  • 25,775+ memory nodes active in production (and counting)
07 /

Horizontal scaling — designed in, not bolted on

Every architectural decision was made knowing this would need to serve thousands of concurrent sessions. The pipeline is stateless per-request; state lives in the DB.

20–120
DB connections
Configurable PostgreSQL pool. Co-located in production → ~1ms reconnect.
5min
Redis cache TTL
Azure Redis caching NPC/player/project profiles. Warm-session pre-load fires after every reply.
5→1
LLM providers
5 today (Azure + Venice + Gemma + Mistral dev). Target: 1 EU-sovereign stack — Mistral 3-model (large / 7B-instruct / 7B-uncensored LoRA) on EU GPU. One provider, zero per-token billing.
7
Parallel groups
Each message fans out across Groups A–D concurrently, then converges at the prompt-build step. No serial bottleneck in data loading.
08 /

Trait taxonomy, romantic styles & the stat engine

Personality isn't stored as a string. It's a computed graph of 56 synonym groups, 162 antonym pairs, 24 romantic styles, and 9 RPG stats — all wired into the pipeline as computation inputs, not descriptors.

TRAIT TAXONOMY

~1,400 lines. Single source of truth.

56 synonym groups · 564 words · 162 antonym pairs · 63 archetype keywords · 6 trauma family groups · 8 communication styles · 4 attachment styles · 5 conflict styles · 24 romantic styles with gate_modifier + stage-based behaviour (early / falling / committed) · 43 consent style aliases · 62 cross-group weighted bridges. Consumed by: unified_psychology, behavior_feedback_service, vad_dynamic_system, dealbreaker_detector.

ARSEAIRSCRAP STATS

9-stat RPG engine wired into psychology

Agility · Resilience · Strength · Endurance · Awareness · Intelligence · Resolve · Sense · Charisma — 1–10 scale, stored in player_attributes and npc_attributes. Used in combat damage math (health_drain_step), strength comparisons in intimate scenes (scene.py), and overwhelm logic. Plus Health 0–200, Mana 0–100, Prestige, Rumours, Advantage. Physical attributes (height, weight, build, appearance, intimate details) in a separate shared table keyed by entity_id + entity_type.

INTIMACY GATES

Personality-driven, not corporate policy

Gates stored in npc_attributes.attraction_preferences JSONB: action_gates · pose_gates · hotspot_gates · progression_gates · hard_limits · forbidden_acts · turn_ons · turn_offs · strictness · session_tolerance · consent_style. IntimacyGate class evaluates every escalation attempt. Boundaries emerge from who the NPC is — not a content filter. A shy NPC with trust 0.10 behaves differently from a confident NPC at trust 0.80 with the same act requested.

HABITUATION COUNTERS

Relationship depth changes what's acceptable

npc_attitudes.metadata.intimate_topic_turns (cap 60) + cooperation_turns (cap 80) — all-time counters written by the background deferred step. familiarity_comfort = clamp((trust+intimacy)/2 + min(ic/100,1)×0.15 + min(coop/40,1)×0.15). At threshold, unlocks "established_partner_rapport" prompt section — softens tone for trusted partners while keeping hard-limits intact. Fixes the case where an NPC with trust=0.79 was still responding like a stranger.

09 /

Why this is impossible with "just vibecoding"

01

It doesn't fit in a context window

No model can hold 1.09M lines in its head. Vibecoding has no memory of decisions made 400 files ago. The system stays consistent because the architecture — not the model — enforces the rules. That architecture had to be designed.

02

The hard parts are invisible in a prompt

Real work lives in edge cases discovered by running the thing against reality: dealbreaker false-positives, cross-NPC awareness, memory-wipe bugs, psyche-drift flush() missing. Vibecoding produces the happy path. These are the 80% that only show up under real use.

03

1,296 endpoints have to agree

An emotion written by one service has to be read correctly by ten others — survive an NPC clone, a memory write, a prompt build, a render. Keeping the contracts between pieces stable across 588 service modules is a design job, not a generation job.

04

The throwaway pile is the proof

1,031 _audit_* / _check_* / _fix_* scripts — 132K lines of pure investigation. Hypotheses, measurements, corrections. Vibecoding skips all of it, which is exactly why vibecoded apps look done but fall apart under use.

10 /

The fossil record — real documented fixes

Each of these is a real fix in the codebase — a case where the obvious implementation was wrong in a subtle way, someone reasoned out why, and re-architected. No prompt generates the insight that finds these. Only running the real system against real users does.

dealbreaker-false-positive-fix cross-npc-awareness-fix hard-limit-fabricated-assault interaction-count-inflation-bug memory-wipe-bug-and-fix chat-lethality-system npc-defensive-slap-system language-pollution-bug training-double-save-bug psyche-drift-flush-missing chronic-motivation-loop-never-flushed location-container-vs-leaf-desync player-fact-extractor-npc-identity-poison prompt-leak-3-layer-defense bilingual-prompt-redundancy interaction-count-inflation-bug
11 /

The intelligence is between the files

✗ Vibecoded

  • Persona card → LLM → response. Done.
  • Layers smeared together; each endpoint is a special case.
  • No stable contracts between 588 services.
  • Memory resets each session or is a context-window hack.
  • NPC "personality" is a system prompt string, not a computed state.
  • Looks done in a demo — falls apart under real sustained use.

● Designed system

  • Layering: routers (what's asked) vs. services (how it's decided) vs. database (what's true). 1,296 endpoints stay consistent because the layer enforces it.
  • Stateful identity over time: every message permanently changes 9 simulation systems. The NPC at message 600 is measurably different from message 1.
  • Autonomy: NPCs have drives, dreams, agendas, and the ability to walk out. Not driven by the player — driven by their own unmet needs.
  • Failure-driven correction: every fix in the fossil record is reasoned judgment. That judgment is the part a model can't supply for you.

Vibecoding can write lines. It cannot hold a million of them consistent, can't discover the edge cases that only show up in real use, can't wire nine simulation systems so they permanently change each other, and can't make an NPC that one day walks out because it has its own life. You didn't write an app. You designed a system — and the design is the work. The 1.09 million lines are just where it landed.

12 /

ELMU — the infrastructure behind it

Virtual Twilight is the consumer stress-test of a B2B/B2G infrastructure product: ELMU — European Living Memory Unit. The same runtime that powers VT is licensed to European institutions as an API. The engineering above is proof it is real.

🏛️

Institutional customers — in production

Named API clients using the ELMU runtime in vocational training and healthcare simulation:

  • Vamia — 10 training NPCs, 131+ goals, 167+ behaviors, 5 healthcare scenarios in production
  • OSAO · TAMK · Karelia UAS · University of Helsinki
  • 3D environments · Real student sessions · Measurable learning outcomes
💰

Cost economics — the real lever

Cost is dominated by large input context (~6.8K–12K tokens/message), not the short reply. The path is self-hosting, not prompting.

  • Azure GPT-4o today: ~€285 / 10K turns
  • Azure GPT-4o-mini today: ~€10 / 10K turns
  • Target: 100% self-hosted Mistral on EU GPU: <€0.02 / 10K
  • ~500× reduction vs GPT-4o-mini. Amortised GPU time, zero per-token billing.
  • Three Mistral models (large / 7B-instruct / 7B-uncensored LoRA) replace all 5 current providers on the same EU bare-metal deployment. Character LoRAs trained on CINECA Leonardo (already in production for fine-tuning).
🇪🇺

EU-sovereign by design

GDPR / EU AI Act compliance is architectural, not retrofitted. At inference time, personal identifiers are replaced with event abstractions — agents reason about world events, not people. Fine-tuning already on CINECA Leonardo (EU HPC). Registered in Finland.

  • No VC · No advisory board · No outside pressure
  • Production inference migration target: OVH / Scaleway EU bare-metal
  • EU AI Act enforcement begins August 2026 — compliant alternative that exists now

✗ US incumbents (Character.AI · Inworld · Replika · Convai)

  • Session-bound or shallow memory — resets every chat
  • Prompt-level emotion or none — permanent drift doesn't exist
  • GDPR compliance bolted on; US data governance; US cloud
  • Per-token pricing — cost scales linearly with every depth improvement
  • Features removed overnight (Replika 2023, Character.AI 2025)
  • No architectural path to EU sovereignty

● ELMU — EU-native persistent agent infrastructure

  • Continuous episodic+semantic DB memory — no ceiling, personality-filtered
  • Continuous VAD with permanent drift, dream-psyche loop, 9-drive motivation
  • GDPR/AI-Act compliance architectural — Event Abstraction Layer at inference
  • Target: <€0.02/10K — cost becomes amortised GPU time, not per-token
  • TRL 7: 56 agents · 6 worlds · 25,775 memory nodes · real revenue · real users
  • EU-based · own models · own infrastructure · bonds are permanent