ReaLMM is a self-hosted LLM operator built on top of LiteLLM router. A harness to run an agent includes:
- Unified LLM gateway with automatic provider detection and OpenAI-compatible endpoints
- Memory
- Policy enforcement and security
- Routing: Retries, cache, and fallbacks
- Observability and tracing
- Operator console
- Python 3.12
- Node.js 22
- Docker and Docker Compose (if you use the published operator stack
gateway+web+ Redis). - At least one model source: a provider
*_API_KEYin.env(Groq, OpenAI, Anthropic, and others LiteLLM can detect), or a running local model using Ollama - Inbound auth:
GATEWAY_ALLOW_OPEN=1is loopback-only. Compose forces it off, so setGATEWAY_API_KEYandGATEWAY_KEY_PEPPERor the gateway will refuse to start.GATEWAY_KEY_PEPPERis also required when any issued key exists, and before a new key is hashed, including open mode. Open mode with an empty key file may still serve the legacy gateway key. - Redis for Compose.
WEB_CONCURRENCYorUVICORN_WORKERSgreater than 1 withoutREDIS_URLrefuses to start.GATEWAY_ALLOW_SPLIT_BUDGET=1restores a warning and continues. Redis covers cache, RPM, and the daily budget. The key file anddata/runtime-flags.jsonstay on one process. A single local process can run without Redis. - Optional extras:
requirements-pii.txtplus a spaCy model ifPII=1; Langfuse keys if you want remote prompts and tracing instead of localprompts/*.json.
There are two install paths. Compose is the published operator stack, a host venv is the local loopback path.
git clone https://github.com/M1ndSmith/ReaLLM.git
cd ReaLLM
cp env/groq.env .envEdit .env and set:
GROQ_API_KEY(or another provider key / Ollama)GATEWAY_API_KEYGATEWAY_KEY_PEPPER
Compose forces GATEWAY_ALLOW_OPEN=0, so the last two are required or the gateway will not start.
docker compose up --build -dGateway: http://127.0.0.1:8000
Console: http://localhost:3000
git clone https://github.com/M1ndSmith/ReaLLM.git
cd ReaLLM
cp env/groq.env .envFill a provider key. Presets set GATEWAY_ALLOW_OPEN=1 for this loopback-only run.
python3.12 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python -m uvicorn app.main:app --host 127.0.0.1 --port 8000Console, in another terminal:
cd web
cp .env.local.example .env.local
npm install
npm run devSet the required values in .env, then:
set -a && . ./.env && set +a
curl -s http://127.0.0.1:8000/healthz
curl -s http://127.0.0.1:8000/models \
-H "Authorization: Bearer ${GATEWAY_API_KEY}"Skip the auth header only if you started the host path with GATEWAY_ALLOW_OPEN=1 and no gateway key.
curl -s http://127.0.0.1:8000/chat \
-H "Authorization: Bearer ${GATEWAY_API_KEY}" \
-H "Content-Type: application/json" \
-d '{"model":"groq/openai/gpt-oss-20b","messages":[{"role":"user","content":"hello"}]}'Use an id from /models. For OpenAI SDKs, set base_url to http://127.0.0.1:8000/v1 and api_key to the gateway key.
Open http://localhost:3000. Paste the gateway key if asked. Playground chats, Connect copies clients, Settings toggles layers and issues keys.
Operator details (scopes, sidecars, errors): USAGE_WALKTHROUGH.md.
.env(dotenvfile, no default): supplies gateway and provider settings to Docker Compose.GROQ_API_KEY(string, no default): required by the includedenv/groq.envprovider preset.GATEWAY_API_KEY(secret string, no default): required because Docker Compose setsGATEWAY_ALLOW_OPEN=0.GATEWAY_KEY_PEPPER(secret string, no default): required when Docker Compose enables gateway authentication.- Optional layers live in
config/realmm.yaml. Host presets:env/selectsconfig/. Env overrides YAML. Memory uses local FastEmbed unlessmemory.embedderisopenaior another embedding model id.nvidia/nemotron-3-embed-1bis one of those ids. Other ids needmemory.embedding_dims.env/nvidia.envpoints atconfig/nvidia.yaml(Nemotron chat, embeddings, and content safety). Stored facts are scoped to the authenticated key (defaultforGATEWAY_API_KEY). A clientuser_idis trace metadata only. Facts already stored underlocaldo not appear under a real key. Guards needguards.enabledplusgroq/meta-llama/llama-prompt-guard-2-22mandgroq/meta-llama/llama-guard-4-12b. With PII on, a streamed reply is buffered and redacted once at the end. The chat estimate is reserved before the provider call. Operator details: Usage walkthrough.
See CONTRIBUTING.md.
Distributed under the MIT License. See LICENSE for more information.
LinkedIn
a.eljoaydi@gmail.com
X
Project Link: https://github.com/M1ndSmith/ReaLLM



