Local-only Ollama workbench with chat, tool calling, PDF/image OCR, invoice extraction, local voice input, and embedded vector memory/RAG.
dotnet run --urls http://127.0.0.1:5288Open http://127.0.0.1:5288.
LLLM stores memory locally in an embedded SQLite file. There is no Docker database or background DB server.
App_Data/memory/lllm-memory.sqlite
Embeddings are generated through local Ollama using nomic-embed-text:latest. Uploaded PDF/image/text documents are chunked, embedded, and searchable from the Local Memory panel. Relevant memory and document chunks are automatically retrieved before each chat turn and shown in the assistant response under Memory Used.
Open LLLM in Chrome, Edge, or Safari and use the browser's install/add-to-dock option. The installed PWA still talks only to the local LLLM server on loopback.
Create a self-contained Apple Silicon binary:
scripts/publish-macos.shRun it directly:
artifacts/lllm-osx-arm64/LLLM --urls http://127.0.0.1:5288The executable is artifacts/lllm-osx-arm64/LLLM.
After publishing, install the macOS LaunchAgent:
scripts/install-launch-agent.shRemove it:
scripts/uninstall-launch-agent.shLogs are written under App_Data/logs.
Check the LaunchAgent status:
launchctl list | grep com.maxneovici.lllmVoice input uses local whisper.cpp; no browser speech API or cloud service is used.
One setup path on Apple Silicon:
brew install whisper-cpp ffmpeg
mkdir -p models
curl -L -o models/ggml-base.en.bin https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.en.binFastest model option:
curl -L -o models/ggml-tiny.en.bin https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-tiny.en.binThen update appsettings.json if needed:
{
"LocalAi": {
"Speech": {
"WhisperExecutable": "/opt/homebrew/bin/whisper-cli",
"ModelPath": "models/ggml-base.en.bin",
"FfmpegExecutable": "/opt/homebrew/bin/ffmpeg"
}
}
}Use the Mic button to record. LLLM transcribes short local chunks into the prompt box and auto-sends after a few seconds of silence.