Launch offer: 40% off for 6 months — just $5.99/mo (reg. $9.99)Claim the deal
Power users

Bring your own AI model

Point your assistant at an AI model running on your own Mac instead of Claude in the cloud. Day-to-day replies then cost nothing per message, and the conversation never leaves your machine. You may see this called BYOLLM — bring your own large language model.

Advanced users only — and here is exactly why

  1. Some models are confidently wrong. In our testing, a less-capable model once reported the result of a step it hadn't actually finished. Whether yours does that depends on the model you pick — so be comfortable checking its work until you know how it behaves.
  2. Some models struggle with multi-step jobs — reading files, running things, handling pictures, which is most of what an assistant does. That is where the less-capable models we tried fell down first. A more capable model, and the hardware to run it, is what closes that gap.
  3. If your computer sleeps, loses its network, or the model stops, your assistant stops answering. It will not quietly switch back to Claude, because that would spend money you chose not to spend.
  4. You choose the model, so you own the answers. We can tell you whether your assistant can reach your computer. We can't support the computer itself, or the model you picked.
  5. Your helper still uses Claude. Dex — the part that maintains your setup and fixes problems — keeps using your Anthropic key. It just costs far less.

What you need

  • An Apple-silicon Mac with enough memory. This is the part people underestimate. A large model needs roughly 15GB to itself, on top of your assistant and macOS — comfortable on a 32GB Mac, and on a 16GB Mac only with a noticeably smaller model.
  • oMLX installed, with a model downloaded and serving. Your assistant reads the list of models from oMLX itself, so there is no model name to type.
  • Your assistant on your own hardware — on that same Mac, or on another computer in the same home. See the two setups below.
  • Your Anthropic key still connected. Dex keeps using it to look after your setup, at a small fraction of what conversations cost.

What it felt like in our testing

Typical reply time, local model
~10 sec

Median across a 22-turn session on an M-series Mac, with a range of 5 to 46 seconds. Your own numbers depend entirely on your machine and the model you choose — a bigger model on a smaller Mac will be slower.

One thing worth knowing before you pick: of the two models we tried, one showed its private reasoning as part of its replies and the other kept it to itself. It is worth a couple of test messages to see which sort you have chosen.

Two setups that work

Your assistant has to be able to reach your Mac. Today that means being on the same machine, or on the same home network.

One Mac, nothing else

Your assistant and the AI model both run on the same Mac. This is the simplest version and the one we recommend: no second computer, no network settings, nothing to configure between them.

Two machines, one home network

Your assistant runs on a Raspberry Pi or spare computer, and reaches the Mac across your own wi-fi. Useful if you want the assistant on something small and quiet that stays on, and the model on your bigger Mac.

Not from anywhere else — not yet.

A cloud server sits in a data centre and has no route into your home, so this option needs your assistant to be running on your own hardware too. Worth knowing before you start: where an assistant lives is set when you create it, so an assistant already running in the cloud can't be moved onto your Mac — you would be setting up a new one on a computer you own. Reaching a model at home from a cloud server means opening a door into your network, and we would rather not ship that until we can do it safely.

What it does to the bill

Conversations become free

Every message you send the assistant is answered by your own machine, so there is no per-message charge. Electricity is the only running cost.

Dex still costs a little

The part that maintains your setup keeps using your Anthropic key. It runs rarely and cheaply compared with day-to-day conversation.

Your subscription is unchanged

This is not a different plan and not an add-on. It is a setting in your dashboard. See pricing.

What happens when it stops

The likeliest thing to go wrong isn't sleep — it's a restart.

You start oMLX by hand, and after your Mac reboots nothing starts it again. Your assistant comes back on its own; the model doesn't. The Mac looks perfectly healthy, because it is — only the thinking is missing. We know this one well: our own test machine sat that way for six days. Build the habit of opening oMLX after a restart, the way you would any other app you rely on.

Computers sleep, reboot and drop off the network. When that happens, you will almost certainly find out before we do.

  1. You notice first. A message you send stops coming back — it gives up after about three minutes.
  2. We spot it within about twelve minutes and email you once to explain what happened. One email, not a stream of them.
  3. Nothing switches back on its own. Your assistant will not quietly start spending your Anthropic credit again to cover the gap. That is your call to make, not ours.
  4. Switching back to Claude is one click in your dashboard, and one click to return once the Mac is awake.

One honest limitation: if the assistant and the model are on the same Mac and that Mac goes to sleep, the part that would have emailed you is asleep too. On a one-machine setup, the failed message is the alert.

Your model, your machine, your call

It is a setting in the dashboard, not a different product — and it is reversible in one click.

Get startedRun it on your own computer

Read the blog →