From Minecraft to Production: A Sovereign AI Reference Architecture for On-Prem LLM Apps
14.8M+ decisions, 85K gradient updates, zero cloud dependency. We built a fully on-prem Minecraft RL network spanning Jetson edge inference, RTX 5090 training, and a DGX Spark LLM oracle—a living blueprint for sovereign physical AI you can steal and adapt.

From Minecraft to Production: A Sovereign AI Reference Architecture for On-Prem LLM Apps
We built a fully on-prem, air-gapped Minecraft reinforcement learning network as a living reference architecture for sovereign AI-driven apps: edge inference, high-throughput training, and an on-prem LLM advisor that improves learning without taking the wheel.
If you want "sovereign AI," you eventually run into a practical constraint: you either own the compute path end-to-end (training, inference, governance, observability), or you outsource your most valuable feedback loops to someone else's cloud.
At Enfuse, we focus on the mechanics. To make it tangible, we built a live, on-prem reinforcement learning network that trains autonomous agents in Minecraft—entirely on local, air-gapped infrastructure. Minecraft is a sandbox, but it's not a toy: it's a clean proxy for the real world—perception, planning, latency, failure modes, and long-horizon goals.
This project is published as a set of living system docs:
- Project Docs: https://mc.enfuse.ai/
- System Overview: https://mc.enfuse.ai/index.html
- DGX Spark Oracle: https://mc.enfuse.ai/dgx.html
- ALA System Architecture: https://mc.enfuse.ai/ala.html
The concept: split the agent's "consciousness" across purpose-built nodes
We are building a superhuman Minecraft agent by splitting its consciousness across three specialized machines. Each tier handles what it's built for—latency-critical perception at the edge, high-throughput training in the middle, and semantic reasoning at the top.
1) The Body — Edge Inference on Jetson Orin AGX
A Jetson Orin AGX "lives in the game." It sees pixels and moves limbs. The system runs 5 Mineflayer bots simultaneously, with the policy network executing inference at 9.5ms latency—fast enough to react to the game world as it changes.
This mirrors how we deploy edge inference for real-world environments: camera feeds, LiDAR, kiosks, robotics, and physical spaces where latency matters. The agent's "body" is always on, always processing, never waiting for a round-trip to a data center.
2) The Brain — Training Hub on RTX 5090
An RTX 5090 host learns from huge batches of experience data using Proximal Policy Optimization (PPO). This is where throughput and iteration speed matter—turning raw experience into improved behavior with repeatable training loops.
The training system runs at approximately 21.4 FPS for gradient updates, processing millions of decisions into actionable policy improvements. The separation from the edge node means training can happen asynchronously—the body keeps acting while the brain keeps learning.
3) The Wisdom — LLM Oracle on DGX Spark
A DGX Spark provides high-level semantic guidance through our Asynchronous LLM Advisor (ALA) system.
Hardware specs:
- GPU: NVIDIA GB10 Blackwell
- Model: GPT-OSS-20B (20 billion parameters via vLLM)
- Throughput: ~240 tokens/sec prompt processing
The DGX Spark serves as the AI Oracle—a compact supercomputer running a 20-billion parameter language model, providing semantic reward shaping, hindsight relabeling, and episode summarization.
The key point: no cloud dependency is required to train, iterate, or reason.
Real metrics from a real system
As of January 6, 2026, the system report shows:
| Metric | Value |
|---|---|
| Total Decisions | 14,884,664+ |
| Gradient Updates | 85,238 |
| Observation Dims | 917 |
| Action Space | 38 Actions |
| Inference Latency | 9.5ms |
| Training FPS | 21.4 |
These numbers are not the point by themselves—the point is that the full learning loop is owned, local, measurable, and reproducible. Every decision is logged. Every inference is tracked. Every gradient update is versioned. You can audit the entire chain.
Neural network architecture
The agent perceives the world through a multimodal observation space of 917 dimensions, combining state information with rich visual features.
Policy Network (MLP)
Input (917) → Linear(256) + ReLU → Linear(256) + ReLU → Output(38 Discrete Actions)
Observation Space Breakdown
State Vector (125 dimensions):
- Original Bot State: 82 dims
- Block Grid 3x3x3: 27 dims
- Entity Types: 5 dims
- Entity Hostility: 5 dims
- Biome/Light/Cliff: 6 dims
Vision Features (792 dimensions):
- RGB MobileNet embeddings: 512 dims
- Color Histogram: 48 dims
- Depth Estimation: 64 dims
- Segmentation features: 128 dims
- Game Knowledge: 40 dims
This multimodal architecture lets the agent combine spatial awareness, visual understanding, and game-specific knowledge into a unified decision-making process.
The differentiator: ALA (Asynchronous LLM Advisor)
Most "LLM + agent" demos hand the controls to the LLM. Voyager and SayCan are impressive, but they're designed for environments where the LLM generates code or directly selects actions. That's fine for prototypes, but it's brittle when you need correctness, stability, and repeatable learning.
Our Asynchronous LLM Advisor (ALA) uses a fundamentally different model:
Unlike Voyager or SayCan, ALA does not let the LLM control actions directly. It acts as a "soft voice in the head" inside the PPO loop.
How it works
ALA is a system where a Large Language Model provides continuous, non-blocking strategic guidance to reinforcement learning agents via a bounded logit perturbation channel.
The math is straightforward:
π(a|s) = softmax(f_θ(s) + β · b(m))
Where:
f_θ(s): policy network logitsb(m): LLM-provided bias vector (exogenous, fixed per step)β: scalar bias strength (decayed over training)
What makes it different
-
The LLM does NOT control actions directly. It provides a bias vector that nudges the policy's probability distribution—a suggestion, not a command.
-
Non-blocking guidance. The LLM runs asynchronously. The agent doesn't wait for the LLM to respond before acting. This preserves real-time performance.
-
On-policy correctness is preserved. We store the biased log-probabilities at rollout time. When training happens, PPO sees the actual distribution the agent sampled from—no off-policy corruption.
-
Bounded perturbation. The bias strength
βis decayed over training. Early on, the LLM's guidance has more influence. As the policy improves, it relies less on the oracle and more on its own learned behavior.
What the Oracle actually does
The DGX Spark's LLM doesn't just observe—it contributes three specific functions:
1. Semantic Reward Shaping
The LLM provides reward bonuses by interpreting the semantic meaning of actions in context.
Example: The agent kills a zombie. The base reward is sparse and delayed. But the LLM can recognize that killing zombies is useful for survival, especially at night, especially when the agent's health is low. It provides an immediate bonus signal that shapes learning toward meaningful behavior.
2. Hindsight Relabeling
When an episode fails, most of that experience is wasted—the agent tried something, it didn't work, and the sparse reward signal says nothing about what was learned.
Hindsight relabeling changes this. The LLM asks: "What did the agent actually achieve?"
Maybe the agent tried to build a house but only placed two blocks before dying. The LLM relabels that episode: "Goal achieved: placed two blocks." Now the agent has positive learning signal from a failed episode. Failures become partial successes.
3. Episode Summarization
Long-horizon episodes are hard to learn from because the credit assignment problem explodes. The LLM compresses episodes into semantic summaries that capture what happened and why it mattered.
This summarization feeds back into the training loop as context for future episodes, building a form of episodic memory without requiring the policy network to learn long-term dependencies from scratch.
Curiosity-driven exploration
Beyond the LLM oracle, the system uses an Intrinsic Curiosity Module (ICM) to encourage exploration. The ICM provides curiosity rewards based on prediction error—when the agent encounters something unexpected, it gets a bonus for exploring.
This creates a dual-guidance system:
- ICM drives low-level exploration (try new things, visit new places)
- ALA provides high-level semantic direction (this action is meaningful, pursue this goal)
Live streaming infrastructure
We stream the training live because transparency is sovereignty too. You can watch the agent learn, fail, and improve in real time.
Hardware Accelerated: The streaming pipeline uses NVENC + Vulkan via Chrome ANGLE for low-latency encoding at 1080p60.
- Twitch: https://www.twitch.tv/enfuseio — Watch live training sessions
- YouTube: https://www.youtube.com/@enfuseio — VOD archive and highlights
The entire capture-to-stream pipeline runs on local hardware:
- Dashboard captured via Puppeteer
- Chrome GPU rendering with Vulkan
- Xvfb virtual display
- FFmpeg + NVENC encoding
- RTMP distribution to Twitch, YouTube, and Owncast
Why Minecraft? Because this is Sim2Real you can watch
Minecraft gives you a fast feedback loop for the hard problems in physical AI:
- Partial observability: The agent only sees what's in front of it. No god's-eye view.
- Long-horizon objectives: "Build a house" takes thousands of steps. "Survive the night" requires planning.
- Mistakes and recovery: Fall into a pit? Get attacked by mobs? The agent must adapt.
- Perception-to-action latency constraints: Act in real time or die.
From there, the enterprise mapping is straightforward:
| Minecraft | Enterprise |
|---|---|
| Game environment | NVIDIA Isaac Sim, Omniverse, or your real facility |
| 5 Mineflayer bots | Thousands of agents running in parallel |
| 917-dim observations | Raw sensor fusion—LiDAR + computer vision on Jetsons |
| Sparse rewards | Business KPIs, quality metrics, safety scores |
Same architecture. Different environment.
The mapping to Enfuse's Sovereign AI Platform
This Minecraft RL network isn't a research project—it's a living reference architecture for the Sovereign AI Platform we ship to enterprises.
Runtime Layer
The three-tier compute topology (edge → training → oracle) maps directly to how we deploy production AI:
- Edge: Jetson or industrial GPUs running inference at the point of action
- Training hub: On-prem GPU clusters running continuous improvement loops
- Oracle layer: DGX Spark or larger systems running LLMs for reasoning, planning, and semantic understanding
Factory Layer
The orchestration around this system—the experiment tracking, weight versioning, metric dashboards, and deployment automation—is the same pattern we productize in the App Factory:
- Templates for RL training loops
- Governance-by-default for model versioning
- Observability stack for real-time metrics
- GitOps workflows for reproducibility
MCP Services
The components in this system map to MCP services:
- VoxSovereign: Speech processing (imagine voice commands to the agent)
- Panopticon: Vision processing (the agent's pixel perception pipeline)
- Model Gateway: LLM routing (the oracle's vLLM backend)
- Cloudberry: Data foundation (training data, episode logs, metrics)
"Sovereign AI" needs a runtime and a factory
If you stop at infrastructure, you end up with expensive hardware that runs a few demos. Sovereign AI becomes real when you can repeatedly ship governed AI apps on your own infrastructure, then scale across teams and departments.
That's the thesis behind the Enfuse platform direction: a runtime + app factory approach for governed, on-prem AI delivery—where the intelligence is portable, secure, and owned by the operator.
This Minecraft RL network exists for one reason: to prove the delivery mechanics, end-to-end, under real constraints.
- The hardware is commodity (Jetson, RTX, DGX Spark).
- The software is open-source compatible.
- The architecture is reproducible.
- The metrics are public.
Watch it live (and steal the blueprint)
The architecture is yours to adapt. The patterns work for any environment where you need:
- Edge inference at real-time latency
- High-throughput training on local hardware
- Semantic reasoning from an on-prem LLM
- Air-gapped operation with zero cloud dependency
What's next
We're continuing to iterate on this system publicly. Upcoming work includes:
- Multi-agent coordination: Multiple Minecraft agents working together
- Hierarchical goal decomposition: LLM-driven task breakdown for complex objectives
- Sim-to-real transfer: Mapping learned behaviors to physical robotics via Isaac Sim
If you want help adapting this reference architecture to your environment—edge sensors, facilities, retail, logistics, smart cities, or industrial workflows—reach out and we'll map it to your constraints.
The point isn't Minecraft. The point is proving that sovereign AI works—end-to-end, on-prem, and accountable.
Related Resources
What Is Sovereign AI?
Complete guide to sovereign AI platforms and data sovereignty
On-Prem LLM Deployment
Hardware requirements, deployment patterns, and implementation methodology
Sovereign Compute
GPU infrastructure management for on-premises AI workloads
Air-Gapped AI Platform
Disconnected AI deployment for classified and secure environments