Short answer: on an Apple Silicon Mac running macOS 15 or later, install oMLX with its official DMG for the menu-bar app or the project’s Homebrew tap for a headless server. Start it on 127.0.0.1:8000, add an API key, download one MLX-format model that leaves memory headroom, then test /v1/models and /v1/chat/completions before connecting another application.
oMLX is an Apache-2.0 project maintained at jundot/omlx. It serves text, vision-language, embedding, reranking and other supported MLX models, manages multiple models with eviction and per-model settings, and exposes an OpenAI-compatible endpoint at http://localhost:8000/v1. It is built for Apple Silicon, not Intel Macs, Windows or Linux.
Choose a stable oMLX release
The latest stable release listed when checked on 10 September 2026 was oMLX 0.6.4, released on 29 August. Use the stable download rather than a release candidate or HEAD build, and record the version actually installed. This guide keeps experimental distributed inference and model-specific acceleration out of the first setup.
Do not copy performance numbers from the release notes into a capacity plan. They describe specific models, quantisations, contexts and hardware. Choose a model by its actual files and intended workload, then measure it locally. The site’s llmfit guide can help shortlist models for your hardware, but its fit estimates also require a real runtime test.
Requirements and resource planning
| Requirement | What to check |
|---|---|
| Mac | Apple Silicon. The project documents M-series support; Intel Macs are outside scope. |
| macOS | macOS 15.0 Sequoia or newer according to the current repository. |
| Python | Only needed for source development; the current source path documents Python 3.11–3.13. The DMG bundles its runtime and Homebrew manages its environment. |
| Memory | Unified memory sufficient for model weights, KV cache, runtime and other applications. Leave several gigabytes for macOS. |
| Disk | Model files plus temporary download space, application files, logs and any SSD KV cache. |
| Network | Needed for installation, updates and model downloads. Optional dashboard web-search providers can also send requests externally if enabled. |
A model fitting on disk does not mean it fits in unified memory. Context length, concurrent requests and caches all add overhead. Start with one small quantised instruct model, the safe memory-guard tier and a modest context. Increase only after Activity Monitor and oMLX remain stable through repeated requests.
Choose the DMG or Homebrew route
| Route | Best for | Important difference |
|---|---|---|
| Official DMG | Menu-bar management, guided model setup and in-app updates | Install from the release assets and drag the app to Applications |
| Official Homebrew tap | Headless service, terminal automation and reproducible upgrades | Installs the CLI/server path; manage it with omlx or brew services |
| Source checkout | Contributors testing code or optional kernels | Needs Git, compatible Python and possibly full Xcode; not the normal beginner path |
This walkthrough uses Homebrew because every server setting can be documented. If you prefer the app, download the DMG from official releases, verify that the publisher and download match the project, drag oMLX into Applications and follow the welcome flow for model directory, server and first model. Do not bypass a Gatekeeper warning blindly or download a similarly named build from a reposting site.
1. Install oMLX with Homebrew
Install Homebrew from its official instructions if it is not already present. Then add the project’s tap and install the formula named in the oMLX README:
brew tap jundot/omlx https://github.com/jundot/omlx
brew install jundot/omlx/omlxConfirm which executable will run and record its version. The exact version command can change, so use the supported help output as the first check:
which omlx
omlx --help
brew info omlxwhich omlx should resolve inside the Homebrew prefix. brew info should show the installed stable version rather than a HEAD build. The optional native-kernel HEAD route in the repository is for specific model families and requires full Xcode; leave it out of a first installation.
2. Create a model directory and start locally
Keep oMLX on the loopback interface until you have a separately reviewed LAN security design. Create the documented default model directory and start a foreground server so errors remain visible. Set a generated API key in your shell rather than copying the example text:
mkdir -p ~/.omlx/models
read -s "OMLX_API_KEY?Choose an oMLX API key: "
export OMLX_API_KEY
omlx serve --host 127.0.0.1 --port 8000 \
--model-dir ~/.omlx/models \
--memory-guard safe \
--api-key "$OMLX_API_KEY"These examples use macOS’s default zsh shell. Keep the server terminal open. Open a second terminal for the following checks; enter the same key again because a separately opened terminal does not inherit the variable from the first:
read -s "OMLX_API_KEY?Enter the same oMLX API key: "
export OMLX_API_KEYConfirm the listener:
lsof -nP -iTCP:8000 -sTCP:LISTENThe server should listen on loopback, not *:8000 or 0.0.0.0:8000. oMLX supports --host, so there is no need to expose it broadly for local clients. A local endpoint still deserves an API key: browsers, plugins and other local processes can reach loopback.
The repository says settings persist in ~/.omlx/settings.json and CLI flags take precedence. Once that file exists after server setup, treat it as sensitive and restrict it to your account:
chmod 600 ~/.omlx/settings.jsonDo not paste the API key into screenshots, Git repositories or support tickets. If another local user can read your home directory or take over the browser session, the key is not a strong isolation boundary.
3. Download one compatible MLX model
Open http://127.0.0.1:8000/admin and authenticate with the key. The project’s current README says the dashboard can download models directly and that oMLX discovers model subdirectories under the configured model directory. Select an MLX-format model supported by mlx-lm or the documented oMLX model family, not a GGUF, GPTQ or generic PyTorch checkpoint.
- Read the model card, licence, intended use and required chat template.
- Check the total download size and available disk before starting.
- Prefer a small quantisation for the first test and leave macOS memory headroom.
- Do not assume a model with the same marketing name uses the same file format.
- Keep gated-model tokens out of the oMLX model directory and screenshots.
After the download, refresh model discovery, select the model and load it. The dashboard can set an API-visible alias. Record the exact ID returned by /v1/models; API clients must use that value or a configured alias, not a name guessed from the Hugging Face page.
4. Validate the OpenAI-compatible API
First confirm that a request without a key is rejected. The README documents a localhost verification-bypass setting in the admin panel; keep that bypass disabled for this setup. In the second terminal, request the catalogue without credentials:
curl -sS -o /dev/null -w "%{http_code}" \
http://127.0.0.1:8000/v1/modelsExpect an authentication rejection such as 401 or 403, not 200. If you get 200, review the admin authentication settings before treating the key as enforced. Then ask for the catalogue with the key. After the model is loaded, check for a successful HTTP status and a JSON model list:
curl -sS -i http://127.0.0.1:8000/v1/models \
-H "Authorization: Bearer $OMLX_API_KEY"Copy the exact id value into a temporary environment variable, then send a low-temperature, low-token validation request. Replace the placeholder rather than leaving it literal:
export OMLX_MODEL_ID='copy-the-exact-model-id-here'
curl -sS http://127.0.0.1:8000/v1/chat/completions \
-H "Authorization: Bearer $OMLX_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<JSON
{"model":"$OMLX_MODEL_ID","messages":[{"role":"user","content":"Reply with exactly: oMLX is ready"}],"temperature":0,"max_tokens":20}
JSONA successful keyed response, together with the rejected unkeyed request, checks routing, key enforcement, model selection and generation. It does not prove the model follows every instruction or that tool calling, structured output, vision and long context work. Test each capability separately with the exact client and model you plan to use.
5. Connect an OpenAI-compatible client
Configure the client with these three values:
- Base URL:
http://127.0.0.1:8000/v1 - API key: the generated oMLX key, supplied through the client’s secret store or environment
- Model: the exact model ID returned by
/v1/models
Do not insert the key into a public browser-based application. A web page can expose it to source code, browser storage or extensions. Desktop clients differ in how they handle a custom base URL, streaming, tool calls and TLS. Validate one plain chat completion before enabling agents or tools that can read files or run commands.
If you need a cross-platform Docker alternative rather than an Apple Silicon-native server, see the site’s LocalAI on Windows guide. Do not install both merely because they expose a similar API; choose the runtime that matches the hardware and operational model.
6. Run oMLX as a managed service
Before switching to the background service, save the API key through oMLX’s admin settings. The foreground --api-key argument is not persisted by the CLI settings save. Keep the option to skip API-key verification disabled. After starting the service, repeat both catalogue requests from section 4: the request without credentials must be rejected, and the request with your key must succeed. Stop the service and fix authentication if either result differs.
After the foreground test succeeds and your saved settings are correct, stop it with Control-C. The current repository documents portable lifecycle commands and Homebrew service commands:
omlx start
omlx stop
omlx restart
brew services info omlxHomebrew service logs are under $(brew --prefix)/var/log/omlx.log, while the structured server log is under ~/.omlx/logs/server.log. Confirm the service still binds to loopback and that the API key is required after restart. Auto-restart is convenient, but it can also restart a memory-heavy model after a crash, so monitor repeated failures rather than accepting a green service state.
Security and privacy boundaries
- Local inference is not automatically offline: installation, model downloads, updates and optional dashboard web search use network services.
- Keep loopback as the default: LAN exposure needs TLS, firewall rules, key rotation and a threat model that this beginner guide does not provide.
- Protect credentials: restrict the settings file, use client secret storage and rotate the key after accidental disclosure.
- Review model licences: oMLX’s Apache licence does not change the licence of a downloaded model.
- Review client tools: an agent connected to a local model can still access files or execute commands if its client grants those tools.
- Plan SSD cache retention: cached prompt state and logs can persist data beyond the request. Configure and clean them according to your sensitivity requirements.
Troubleshooting oMLX
| Symptom | What to check |
|---|---|
omlx: command not found | Run brew info omlx and which brew. Add the correct Apple Silicon Homebrew prefix to the shell path; do not install a second unrelated package. |
| Port 8000 is busy | Use lsof to identify the owner. Stop the intended service or choose another local port and update the client base URL. |
No models in /v1/models | Confirm the model is MLX-format, sits under the configured directory, has completed downloading and is loaded. Check the server log. |
| 401 or 403 response | Confirm the same API key reaches the server, without extra quotes or whitespace. Rotate the key if its handling is uncertain. |
| Model loads, then macOS kills the process | Use a smaller quantisation, safe memory guard, shorter context and fewer concurrent requests. Unload pinned models and leave more system headroom. |
| Chat response is empty or malformed | Check that the model has a compatible chat template and use the exact model ID. Test plain non-streaming chat before tools or reasoning options. |
| Client cannot connect | Verify curl works on the same Mac, then inspect client proxy settings, base URL, API path and whether it requires HTTPS. |
Update, rollback and remove oMLX
Before updating, record the installed version, stop active clients, back up only the settings you understand and confirm model directories are not the sole copy of valuable work. Upgrade the Homebrew installation with:
omlx stop
brew update
brew upgrade omlx
omlx startRepeat the listener, authentication, model-list and chat tests. If an update breaks your workflow, stop the service and consult the official release notes before installing an older release. Rolling back binaries while retaining a newer settings schema can create a second problem, so preserve a versioned settings backup.
To remove the Homebrew package, stop it and use brew uninstall omlx. Untap jundot/omlx only if you do not use another formula from that tap. Uninstalling the package might leave ~/.omlx, downloaded models, logs and caches. Inspect their exact sizes and contents, then move unwanted data to the Trash manually; do not issue a broad recursive delete against your home directory. For the DMG route, quit the app and move oMLX.app to the Trash, then review the same data locations.
Frequently asked questions
Does oMLX work on an Intel Mac?
No supported Intel route is documented by the project. oMLX is built around Apple’s MLX stack and requires Apple Silicon.
Can oMLX load GGUF models?
This guide uses MLX-format model directories. GGUF is a different format generally served by llama.cpp-family runtimes. Download a compatible MLX conversion or choose a GGUF runtime instead of renaming files.
Is the oMLX API completely compatible with OpenAI?
It implements OpenAI-compatible endpoints, but compatibility is feature-specific. Test the endpoint, model and client combination you need. Tool calls, reasoning, images, embeddings and streaming should each be validated separately.
Can another computer use my oMLX server?
Technically the server has configurable bind settings, but this guide deliberately stays on loopback. LAN service needs TLS, authentication, firewalling, network trust and update procedures. Do not switch to 0.0.0.0 merely to make a connection error disappear.
Bottom line
The safest oMLX starting point is one small MLX model, a loopback-only server, an API key and a repeatable curl test. Use the DMG when menu-bar management matters and Homebrew when service automation matters. Do not claim speed or fit until the exact model, context and client have survived your own memory and compatibility tests.
Primary project sources
Sources reviewed: 10 September 2026. Reviewed against the official oMLX documentation and stable 0.6.4 release. Record the application and model versions, then validate the workflow in your own environment.