A Mac mini can run a local LLM for event-driven forex trading without sending prompts to a hosted model provider. This guide compares current open-source and open-weight models, Mac mini memory tiers, five model-download routes, and local inference costs against ChatGPT and Claude subscriptions plus OpenAI, Anthropic, and DeepSeek APIs. The safe design is a local-inference trading service: FXMacroData delivers source-backed release data, the Mac builds a compact evidence snapshot and runs the model, deterministic code applies risk policy, and a broker API handles executable prices and orders.
Best fit
Use it for
Event-driven strategies that can wait seconds, not microseconds, and benefit from private local inference.
Start with
Historical replay, shadow signals, and a broker practice account before any small, tightly capped live trial.
Do not use it for
High-frequency trading, guaranteed-return claims, or unconstrained model-to-broker access.
Cost checkpoint
Compare the same prompt, output, latency target, and replay score across local and hosted models before choosing on price.
What “offline model” means
Offline describes the inference path. Once the model weights, tokenizer, Python packages, and strategy configuration are installed, the prompt and response can stay on the Mac mini. A live system still needs outbound access to receive FXMacroData events and reach a regulated broker. If either connection is unavailable, the correct action is FLAT and a clear alert—not a guess from cached model knowledge.
| Can stay local | Must remain connected |
|---|---|
| Model weights, tokenizer, prompt templates, feature construction, policy rules, audit logs, and historical replays. | FXMacroData release delivery, broker bid/ask, tradeable status, account state, order submission, and transaction reconciliation. |
| The model's proposed direction, horizon, confidence rank, evidence identifiers, and short rationale. | Any fact that can change after the model was downloaded: actual releases, forecasts, revisions, prices, spreads, balances, and open positions. |
This boundary provides useful privacy without pretending a connected market can be traded offline. It also makes failures easier to reason about: a model problem, a macro-data problem, and a broker problem are three separate states.
Mac mini memory: what model size fits?
As of September 2026, Apple lists Mac mini configurations with M6 and M5 Pro. The listed unified-memory ranges are 16–32GB for M6 and 24–64GB for M5 Pro; Apple gave first availability as September 22, 2026. Because the new configurations are so recent, independent benchmarks for this complete trading workload are not yet available. Buy for memory headroom and measure your own event-to-order path.
Apple's MLX research on a 24GB M5 MacBook reported inference-memory demand of about 5.61GB for a 4-bit 8B model, 9.16GB for a 4-bit 14B model, and 17.31GB for a 4-bit 30B-A3B mixture-of-experts model. The benchmark used a 4,096-token prompt and generated another 128 tokens. These are measured inference-workload footprints, not model-file sizes or guarantees for total application memory. The operating system, context cache, Python process, event service, and broker adapter all need room too.
| Unified memory | Practical starting point | What to validate |
|---|---|---|
| 16GB | One small 3B–8B 4-bit model, short prompts, one strategy process. | Memory pressure, context growth, cold-start time, and swap during release bursts. |
| 24GB or 32GB | An 8B or 14B 4-bit model with more context and service headroom. | Prefill time at the exact evidence length and sustained decode while other services run. |
| 48GB or 64GB | Larger quantized models, longer contexts, or isolated parallel services. | Whether the extra model quality improves out-of-sample decisions enough to justify slower prefill. |
For this use case, a smaller, repeatable model often beats a larger, more eloquent one. The prompt should contain a small structured snapshot, and the output should contain one machine-validated proposal. A 3B–14B quantized instruction model is a sensible first measurement range; it is not a promise that any model can forecast FX profitably.
Best open-source and open-weight models for a Mac mini
“Open-source model” is a popular search term, but open-weight is often the accurate description. In this guide, an open model is a checkpoint whose weights can be downloaded and run locally. That does not mean every release satisfies the Open Source Initiative's Open Source AI Definition, or that every licence permits every commercial use. Check the exact model card, base-model lineage, licence, acceptable-use policy, and quantizer before deployment.
The table is a September 2026 engineering shortlist, not a profitability ranking. Memory tiers assume a 4-bit text model, a bounded 4K–8K working context, one active model, and headroom for macOS and the trading services. Advertised maximum context is not a sensible live allocation by default.
| Model family | Licence posture | Practical Mac mini tier | Why test it | Important limit |
|---|---|---|---|---|
| Qwen3 4B / Qwen3.5 9B | Apache 2.0 for these checkpoints. | 4B on 16GB; 9B on 24GB or 32GB. | Compact instruction models with long-context capability and a non-thinking path suitable for bounded structured proposals. | Do not allocate the advertised 262K context just because it exists; the KV cache consumes memory and extends prefill time. |
| Gemma 4 E2B, E4B, and 12B | Apache 2.0 for Gemma 4; older Gemma releases can use different terms. | E2B/E4B on 16GB; 12B on 24GB or 32GB. | Small publisher checkpoints with documented Q4 memory estimates and useful text-instruction variants. | Select the text path deliberately; multimodal weights and image components add memory the FX proposal does not need. |
| Llama 3.2 3B Instruct | Custom Llama Community License and acceptable-use policy. | 16GB baseline. | Small, widely packaged checkpoint that is convenient for testing MLX, Ollama, GGUF, and prompt-format portability. | Open-weight, not permissively open-source; attribution and scale provisions apply. |
| Ministral 3 3B, 8B, and 14B | Apache 2.0. | 3B/8B on 16GB; 14B on 24GB or 32GB. | Official GGUF releases plus documented JSON and function-calling support make it a useful schema-compliance candidate. | Model features never replace output validation or the deterministic risk gate. |
| DeepSeek-R1 Distill Qwen 7B, 14B, and 32B | MIT for the DeepSeek release and Apache 2.0 Qwen lineage; Llama-derived distills retain Llama terms. | 7B on 16GB; 14B on 24GB/32GB; 32B on 64GB. | Useful for comparing a local reasoning model with the separate DeepSeek hosted API. | Reasoning length can vary. Start in shadow mode, impose a hard token and time budget, and pin an explicit checkpoint instead of a mutable latest tag. |
| gpt-oss-20b | Apache 2.0 plus the gpt-oss usage policy; described by OpenAI as open-weight. | Officially supports 16GB systems; use 24GB or 32GB for practical trading-stack headroom. | 21B total parameters but 3.6B active per token, a 128K context capability, and broad local-runtime support. | The complete 12.8GiB checkpoint must still be stored and loaded. The 120B release targets roughly 80GB and does not fit within the 64GB ceiling of the Mac mini configurations covered here. |
| Qwen3.6 27B or a 32B distill | Apache 2.0 for Qwen3.6; verify the exact 32B checkpoint lineage. | 64GB research tier. | A larger challenger for offline replay when smaller models fail the evidence and schema tests. | More parameters can mean slower prefill and a missed event deadline. It should not be the default real-time proposer without measured benefit. |
Choose the winner on your own point-in-time evaluation: JSON-schema pass rate, evidence-ID accuracy, unsupported-claim rate, compliance with FLAT instructions, time to first token, total generation time, memory pressure, and cost per accepted proposal. Parameter count and general benchmark rank are not substitutes for that scorecard.
Ollama vs MLX-LM vs llama.cpp vs LM Studio
A model is the checkpoint; a runtime is the software that loads it. The same family can behave differently after quantization, conversion, prompt templating, or runtime upgrades, so benchmark the complete pinned combination.
| Runtime | Why choose it | Operating note |
|---|---|---|
| MLX-LM | Native Python path for Apple silicon, quantization, streaming generation, and prompt caching. | Prefer the in-process Python API. The project labels its example HTTP server unsuitable for production. |
| Ollama | A fast route to a local prototype and a familiar localhost API. | It binds to loopback by default; set OLLAMA_NO_CLOUD=1 when the host must remain local-only. |
| llama.cpp | Portable GGUF models and a compact native runtime. | Metal support is enabled by default on macOS; benchmark the exact build and model you intend to pin. |
| LM Studio | A visual catalogue plus the lms CLI, memory estimation, MLX/GGUF choices, and a localhost server. |
Good for interactive evaluation and format comparison. App/runtime licensing and each model's licence remain separate questions. |
Mac mini local AI cost vs ChatGPT, Claude, and DeepSeek
The first cost distinction is the most important: a consumer chat subscription is not an unattended trading API. ChatGPT Plus/Pro and Claude Pro/Max are interactive products. Automated model calls require separate developer API billing. DeepSeek's free web chat and its metered API are also separate products.
| Option | Current public price* | Suitable for automated inference? |
|---|---|---|
| Mac mini | M6 starts at US$899*; the M6 24GB/256GB configuration is US$1,099*; M5 Pro starts at US$1,699*. | Yes, after a local model and service are installed. There is no per-token model fee, but hardware, electricity, storage, resilience, and maintenance are real costs. |
| ChatGPT | Plus US$20/month*; Pro tiers from US$100/month* to US$200/month*. | No. Use the separately billed OpenAI API for automation; do not treat a ChatGPT subscription as trading-bot API capacity. |
| Claude | Pro US$20/month*; Max US$100/month* or US$200/month*. | No. Anthropic states that Claude subscriptions do not include Claude API or Console usage. |
| DeepSeek web chat | Free web chat*; API use is metered separately. | No. Use the separate DeepSeek API, or download a compatible DeepSeek open-weight checkpoint and run it locally. |
* Public US-dollar prices retrieved from official provider pages on September 10, 2026. Taxes, regions, promotions, limits, and model availability can change.
Hosted API cost under one transparent workload
The comparison below assumes each eligible FX event sends 4,000 uncached input tokens and produces 300 total output tokens, including any billed reasoning tokens. Low volume is 3,000 decisions per 30-day month (100/day), medium is 15,000 (500/day), and high is 60,000 (2,000/day). It excludes FXMacroData and broker costs because both local and hosted designs need those services.
monthly model cost = input millions × input rate + output millions × output rate
| Hosted API model | Input / output per 1M tokens* | 3,000 decisions* | 15,000 decisions* | 60,000 decisions* |
|---|---|---|---|---|
| OpenAI GPT-5.6 Luna | US$0.20 / US$1.20 | US$3.48 | US$17.40 | US$69.60 |
| OpenAI GPT-5.6 Terra | US$2 / US$12 | US$34.80 | US$174 | US$696 |
| OpenAI GPT-5.6 Sol | US$4 / US$20 | US$66 | US$330 | US$1,320 |
| OpenAI GPT-6 Astra | US$10 / US$50 | US$165 | US$825 | US$3,300 |
| Claude Haiku 4.5 | US$1 / US$5 | US$16.50 | US$82.50 | US$330 |
| Claude Sonnet 5 | US$2 / US$10 | US$33 | US$165 | US$660 |
| Claude Opus 5 | US$5 / US$25 | US$82.50 | US$412.50 | US$1,650 |
| DeepSeek V4 Flash off-peak | US$0.22 / US$0.66 | US$3.23 | US$16.17 | US$64.68 |
| DeepSeek V4 Flash peak | US$0.44 / US$1.32 | US$6.47 | US$32.34 | US$129.36 |
| DeepSeek V4 Pro off-peak | US$0.66 / US$1.98 | US$9.70 | US$48.51 | US$194.04 |
| DeepSeek V4 Pro peak | US$1.32 / US$3.96 | US$19.40 | US$97.02 | US$388.08 |
* Standard API rates retrieved September 10, 2026 from OpenAI, Anthropic, and DeepSeek. The scenario deliberately assumes no cache discount. DeepSeek peak windows are 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday; a market release cannot be delayed merely to obtain an off-peak rate.
Caching can lower repeated-prefix input charges, but only when the provider recognizes a stable prefix, and cache writes can have their own rate. Output is not discounted. Thinking also changes the economics: Anthropic and DeepSeek bill reasoning as output even when the complete chain is not displayed. In the medium scenario, an extra 1,000 output tokens per decision adds about US$9.90 per month for off-peak DeepSeek Flash, US$150 for Claude Sonnet 5, or US$750 for GPT-6 Astra at these rates.
What a local Mac mini costs per month
Illustrative M6 24GB total: about US$28.30 per month
Hardware: US$1,099 ÷ 48 months = US$22.90. Electricity: an illustrative 50W average × 24 × 30 at US$0.15/kWh = US$5.40. Total: US$28.30, excluding taxes, storage, a UPS, downtime, monitoring, and engineering time, with no credit for residual resale value.
Apple documents a 155W maximum continuous-power ceiling and 3W display-on idle consumption, but it does not publish sustained LLM-inference draw. Measure average wall power for the actual workload; 50W here is an explicit scenario, not an Apple benchmark.
At that illustrative US$28.30 monthly local cost and the uncached prompt above, simple price break-even is roughly 813 decisions/day against GPT-5.6 Luna, 86/day against Claude Sonnet 5, 43/day against GPT-5.6 Sol, and 875/day off-peak or 437/day at peak against DeepSeek V4 Flash. Those numbers compare billing only. A cheaper model that fails the same replay, schema, deadline, and risk tests is not an equivalent substitute.
Hosted API wins when
Volume is low or bursty, the best hosted model materially outperforms local candidates, and you prefer variable spend to maintaining model infrastructure.
Local inference wins when
The local model clears the same evaluation, prompt privacy matters, usage is sustained, and you can manage the Mac, storage, power, monitoring, and fail-closed controls.
A hybrid often wins first
Run the small local model on every eligible event; send a limited shadow or escalation sample to a hosted model. Neither path may bypass deterministic policy.
Use a model-proposes, code-disposes architecture
01 · Event
FXMacroData SSE or the changes cursor notifies the strategy when a new or revised release is available.
02 · Evidence
Code fetches the actual and prior values, forecast category, release metadata, and event ID.
03 · Proposal
The local model returns strict JSON: long, short, or flat—never an order.
04 · Gate
Deterministic policy checks data, quote, spread, exposure, loss limits, and duplicates.
05 · Execute
A broker adapter sizes and submits one approved ticket with an idempotency key.
06 · Reconcile
The broker transaction stream confirms, rejects, fills, or closes the ticket.
The safety boundary is architectural, not a line in the prompt. Run the model in a process that cannot read the broker token or call the broker order endpoint. Let it emit only a proposal. A separate execution process owns position sizing, protective orders, daily loss limits, the kill switch, and the broker credential.
Define real-time correctly
This is event-driven trading, not high-frequency trading. FXMacroData emits a release event when a new announcement is available. The Mac then retrieves the corresponding release, constructs features, runs inference, applies policy, obtains a current broker quote, and submits an approved order. The useful latency unit is normally seconds, and the strategy must be profitable after the spread, slippage, and every step in that chain.
Broker streams have their own delivery rules. For example, OANDA documents at most four pricing updates per second per instrument, with the final price from each 250ms window rather than every price created. Treat the broker's bid/ask and tradeable flag as execution truth. FXMacroData FX series are reference context, not executable quotes.
The bars identify measured segments; equal visual length does not imply equal duration. Record real distributions from your machine, model, location, and broker.
Prerequisites
- A Mac mini with enough unified memory for the pinned model plus operating headroom.
- Python 3.10 or newer and one local runtime: MLX-LM, Ollama, llama.cpp, or LM Studio.
- An active FXMacroData API key from API Management.
- A broker practice or demo account with streaming prices, account state, order submission, and transaction history.
- A written strategy specification: allowed pairs, permitted event families, maximum exposure, quote-age and spread rules, loss limits, trading hours, and shutdown behavior.
- A fixed model, tokenizer, prompt, policy, and strategy version for every replay and decision log.
Secure the Mac before running the strategy
- Use a dedicated non-administrator macOS account, disk encryption, automatic security updates, and UTC timestamps.
- Bind local model services to loopback and block inbound internet access.
- Keep FXMacroData and broker secrets out of prompts, model-readable files, and logs.
- Use
launchdfor restartable services; after a restart, reload current broker and event state instead of relying on pre-restart assumptions. - Provide a physical and software kill switch that prevents new orders while preserving reconciliation.
Step 1: Download and pin a local LLM on the Mac mini
Every method needs internet access for the initial download. Once the checkpoint, tokenizer, runtime, and dependencies are present, inference can stay local. The right method depends on whether you value the shortest setup, Apple-native Python, a portable GGUF file, a reproducible publisher snapshot, or a visual model browser.
| Download route | Best for | What lands on the Mac | Main caution |
|---|---|---|---|
| Ollama | Quick command-line setup with a local API. | An Ollama-managed quantized model and manifest. | Use an explicit size tag, disable cloud features, and record the resolved artifact instead of relying on latest. |
| MLX-LM / MLX Community | Python-native Apple silicon inference and local conversion. | MLX weights, tokenizer, and configuration. | Community conversions are not publisher originals; verify the base model and conversion revision. |
| Hugging Face CLI | Exact revision pinning, dry-run size checks, and publisher files. | A cached or named local snapshot in safetensors, GGUF, or another published format. | A download tool is not an inference runtime; gated repositories require acceptance of their terms. |
| llama.cpp / GGUF | Portable quantized files and a compact native server. | One or more GGUF files plus the llama.cpp runtime. | Verify who produced the quantization and keep its exact filename, hash, and source revision. |
| LM Studio | Visual comparison, memory estimation, and MLX/GGUF discovery. | Models in the LM Studio directory plus its local runtime. | Good for evaluation; still pin the selected file and keep the live server on loopback. |
Option A: download with Ollama
Install the signed Ollama macOS app, choose an explicit model size, and download it before the trading window. This is the simplest prototype path.
ollama pull qwen3.5:4b
ollama run qwen3.5:4b
# Start a local-only service when Ollama is not already running.
OLLAMA_NO_CLOUD=1 ollama serve
Ollama binds to 127.0.0.1:11434 by default. Preload the pinned model or configure a suitable keep_alive value so inference on a scheduled release is not delayed by a cold model load.
Option B: download an Apple-native MLX model
MLX-LM can download a compatible repository on first use, or convert a publisher checkpoint into a named local directory. Use mlx-vlm rather than mlx-lm when a chosen checkpoint requires a multimodal runtime.
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install mlx-lm
mlx_lm.chat \
--model mlx-community/Qwen3-4B-Instruct-2507-4bit
mlx_lm.convert \
--model Qwen/Qwen3-4B-Instruct-2507 \
-q --q-bits 4 \
--mlx-path /Users/you/models/qwen3-4b-2507-4bit
Option C: pin a Hugging Face snapshot
The Hugging Face hf CLI is a straightforward route when repeatability matters. First use --dry-run to see the download size. Then replace COMMIT_SHA with the immutable revision shown by the publisher repository.
python3 -m pip install -U huggingface_hub
hf download Qwen/Qwen3-4B-Instruct-2507 --dry-run
hf download Qwen/Qwen3-4B-Instruct-2507 \
--revision COMMIT_SHA \
--local-dir /Users/you/models/qwen3-4b-2507
hf cache verify Qwen/Qwen3-4B-Instruct-2507 \
--revision COMMIT_SHA
This route can also fetch selected files with --include. It is useful for keeping a publisher snapshot separate from the MLX or GGUF conversion used at runtime.
Option D: download GGUF and run llama.cpp
GGUF is the portable format commonly used by llama.cpp and several desktop runtimes. A quick smoke test can pull from a repository, but a live evaluation should download one immutable revision and serve the absolute local path.
brew install llama.cpp
hf download mistralai/Ministral-3-8B-Instruct-2512-GGUF \
--revision COMMIT_SHA --include "*Q4_K_M.gguf" \
--local-dir /Users/you/models/ministral-3-8b
llama-server -m /Users/you/models/ministral-3-8b/FILE-Q4_K_M.gguf \
--host 127.0.0.1 --port 8080
llama.cpp can constrain generation with a JSON schema or grammar. Keep that constraint, but still parse and validate the result in ordinary code.
Option E: download with LM Studio
LM Studio provides a graphical model catalogue, while the bundled lms command can search by format, select a quantization, estimate memory, and start a local server.
lms get --mlx
lms get llama-3.1-8b@q4_k_m
lms load --estimate-only llama-3.1-8b --context-length 8192
lms load llama-3.1-8b --context-length 8192
lms server start
Meta also offers a publisher-direct route through its llama-models package. It requires licence acceptance, and its generated download links expire, so it is most useful when provenance from the original publisher matters more than a one-command desktop workflow.
Record these details before replay testing
- Publisher, repository, exact revision SHA, base lineage, licence, and acceptable-use policy.
- File list, checksum, quantizer, bit depth, weight format, and required disk space.
- Tokenizer, chat template, context limit, runtime version, Python lock, and macOS version.
- A local smoke-test result for strict JSON, maximum output length, deadline, and memory pressure.
- An offline restart test proving the service does not silently fetch a model or call a hosted endpoint.
Prefer safetensors or inspected GGUF/MLX artifacts. Do not enable arbitrary repository code in a live service merely to make a checkpoint load. If a model requires a setting such as trust_remote_code, inspect, pin, and isolate that code before evaluation. Do not upgrade model weights, the runtime, or macOS during an active test window: a “same model name” download can still change.
Step 2: Connect FXMacroData
Use FXMacroData MCP while developing the research workflow: the local assistant can inspect the catalogue, release calendar, announcements, predictions, FX context, and market sessions as tools. Use narrow REST or SSE calls inside the live service so each request is explicit, testable, and easy to replay. A useful first scope is one event and one pair, such as US CPI and EUR/USD.
| Surface | Who uses it | Role in this workflow |
|---|---|---|
| MCP | A local assistant during research, development, and post-trade review. | Discovers available tools and answers bounded evidence questions; it does not route orders. |
| REST | The deterministic event and execution services. | Fetches known calendar, announcement, prediction, FX, and session resources through explicit paths. |
| SSE or changes cursor | The always-on event listener. | Signals that a new or revised announcement is available, then passes its event ID to the evidence service. |
{
"servers": {
"FXMacroData": {
"type": "http",
"url": "https://mcp.fxmacrodata.com",
"headers": {"Authorization": "Bearer YOUR_API_KEY"}
}
}
}
A compatible MCP host discovers tool schemas from the server. A calendar request has the same essential shape as this abbreviated contract:
{
"name": "release_calendar",
"inputSchema": {
"type": "object",
"properties": {
"currency": {"type": "string"},
"start_date": {"type": "string"},
"end_date": {"type": "string"}
},
"required": ["currency"]
}
}
Example prompt for a local MCP client:
Use FXMacroData tools for EUR/USD event research.
Check the next USD and EUR releases, then fetch the latest
available observation for the triggered indicator. Keep market
consensus, official forecasts, nowcasts, and FXMacroData
predictions in separate labelled fields. Return evidence only;
do not create an order or invent a missing value.
That forecast separation matters. “Consensus” is not a synonym for an FXMacroData-generated prediction, a central-bank projection, an official survey, or a nowcast. Preserve the category and provenance carried by the API instead of blending them into one model input.
Step 3: Trigger on live macro events
The FXMacroData API exposes a long-lived Server-Sent Events stream at /v1/stream/events. It emits an event when a new economic release becomes available. Filter the stream to the currencies and indicators the strategy has explicitly approved.
curl -N -H "X-API-Key: YOUR_API_KEY" \
"https://api.fxmacrodata.com/v1/stream/events?currencies=usd,eur&indicators=inflation,policy_rate&payload=compact"
Persist the last event_id. On reconnect, send it as Last-Event-ID so buffered events can be replayed. A service that cannot hold SSE can poll /v1/announcements/changes with its returned next_cursor. If the cursor has fallen outside retention, reload the current state from the relevant /v1/announcements/{currency}/latest endpoint before resuming.
An event is a trigger, not a complete strategy input. Fetch the matching announcement after receipt and verify currency, indicator, units, dates, provenance, and data-quality fields. Dedupe by event ID before inference and again before execution.
Step 4: Build a compact inference snapshot
Transform live responses into a small, typed object before they reach the model. Include only evidence fields used during backtests and replay. Keep the broker quote separate from the FXMacroData macro record, and omit account secrets and available margin entirely.
{
"strategy_version": "eurusd-cpi-v1",
"event_id": "FXMD_EVENT_ID",
"as_of": "DECISION_TIMESTAMP",
"instrument": "EUR_USD",
"macro": {
"announcement_id": "FXMD_ANNOUNCEMENT_ID",
"observation_id": "FXMD_OBSERVATION_ID",
"actual": "API_VALUE",
"prior": "API_VALUE",
"forecast": {
"value": "API_VALUE",
"category": "market_consensus",
"source": "API_SOURCE",
"evidence_id": "FXMD_FORECAST_EVIDENCE_ID",
"generated_at": "PRE_RELEASE_API_TIMESTAMP"
},
"unit": "API_UNIT",
"released_at": "API_TIMESTAMP"
},
"broker_quote": {
"bid": "BROKER_BID",
"ask": "BROKER_ASK",
"time": "BROKER_TIMESTAMP"
}
}
The model output should expose fewer fields than the risk engine uses. It may propose LONG, SHORT, or FLAT, a fixed instrument, a bounded horizon, an evidence list, and a confidence rank. It must not set units, leverage, stop distance, take-profit distance, or an order type.
import json
from mlx_lm import generate, load
MODEL_PATH = "/Users/you/models/pinned-fx-instruct"
model, tokenizer = load(MODEL_PATH)
messages = [
{"role": "system", "content": "Return only the required proposal JSON."},
{"role": "user", "content": build_prompt(evidence_snapshot)},
]
prompt = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
)
raw = generate(model, tokenizer, prompt=prompt, max_tokens=160)
try:
proposal = json.loads(raw)
except json.JSONDecodeError:
proposal = {"action": "FLAT", "reason": "invalid_json"}
required = {"action", "instrument", "event_id", "evidence_ids"}
if not required.issubset(proposal):
proposal = {"action": "FLAT", "reason": "missing_fields"}
if proposal.get("action") not in {"LONG", "SHORT", "FLAT"}:
proposal = {"action": "FLAT", "reason": "invalid_action"}
Use a JSON Schema or typed model in the real service, not just these abbreviated checks. A timeout, extra prose, unknown evidence ID, unsupported pair, or schema error must resolve to no trade.
Step 5: Put deterministic risk in front of execution
The risk gate is ordinary code with fixed configuration. The model cannot rewrite it. Run the gate against freshly fetched broker and account state after inference, because the quote and exposure may have changed while the model was generating.
Every condition must pass
- Known, matching, unprocessed event ID
- Canonical actual and expected unit
- Forecast category and provenance preserved
- Model, prompt, policy, and strategy versions approved
- Strict proposal schema valid
- Fresh, tradeable broker bid/ask
- Spread and slippage limits satisfied
- Position, currency, leverage, and notional caps satisfied
- Daily loss and drawdown limits satisfied
- Protective orders and kill switch available
def approve(proposal, event, quote, account, policy):
checks = [
valid_schema(proposal),
known_unprocessed_event(proposal, event),
canonical_macro_fields(event),
fresh_tradeable_quote(quote, policy),
spread_within_limit(quote, policy),
exposure_within_limits(account, policy),
losses_within_limits(account, policy),
protective_orders_available(),
not kill_switch_active(),
]
return all(checks)
gate_allowed = approve(proposal, event, quote, account, POLICY)
if not gate_allowed:
proposal = {"action": "FLAT", "reason": "risk_gate"}
Model confidence is not a probability of profit. At most, treat it as a ranking feature that has to earn its meaning through out-of-sample calibration. The risk engine computes size from account equity, stop policy, exposure, and a hard maximum; it never copies a size proposed by the model.
Step 6: Submit and reconcile through a broker adapter
Use a broker adapter with separate practice and live configurations. The adapter should expose a narrow interface such as get_quote, get_account_state, submit_order, and reconcile. The model process should not be able to import or call it.
event_id = evidence_snapshot["event_id"]
client_id = f"{event_id}:{evidence_snapshot['strategy_version']}"
if gate_allowed and proposal["action"] != "FLAT":
ticket = risk_engine.build_ticket(
proposal=proposal,
quote=broker.get_quote("EUR_USD"),
account=broker.get_account_state(),
client_id=client_id,
)
broker.submit_practice_order(ticket)
broker.reconcile(client_id)
The client ID makes retry behavior explicit. If the network fails after submission but before acknowledgement, query orders and transactions for that ID before retrying. Never assume “no response” means “no order.” Follow the broker's documented recovery model; OANDA, for example, recommends an initial complete account snapshot followed by account updates keyed by transaction ID.
Protective stop and take-profit instructions should be created atomically where the broker supports it. Record the source event, evidence hash, model and policy versions, pre-trade quote, risk result, broker request, broker response, fills, rejects, and later position state. The detailed OANDA v20 and FXMacroData guide shows how the broker and macro layers stay separate.
Step 7: Replay, measure, and stage the rollout
Do not begin with live money. A local model can produce syntactically valid but economically weak answers, and sudden market changes remain unpredictable. The CFTC warns that AI trading bots cannot predict the future or sudden market changes. Treat this as experimental automation with leveraged-market risk.
1 · Offline replay
Use point-in-time data, realistic spreads and slippage, fixed versions, and untouched holdout periods.
2 · Shadow signals
Run against live events and quotes, but create no broker order. Compare intended and feasible fills.
3 · Practice account
Exercise submission, duplicate prevention, protective orders, reconnects, rejects, and reconciliation.
4 · Tightly funded live
Only after written acceptance criteria, independent review, strict position and loss limits, and a rehearsed shutdown path.
A replay must use the information actually available at the decision timestamp. The point-in-time backtesting guide explains why revised macro data cannot be substituted into an earlier decision. Test the whole pipeline, including costs and rejected orders, rather than judging the language model in isolation.
Evaluate local and hosted models against the same criteria
| Evaluation metric | What it reveals | Example acceptance criterion |
|---|---|---|
| JSON-schema pass rate | Whether the runtime, quantization, and prompt reliably produce a parseable proposal. | Meet the written minimum on the full holdout set; invalid output always becomes FLAT. |
| Evidence-ID accuracy | Whether each claim maps to an observation, announcement, forecast, or quote that was actually supplied. | No invented identifier may reach the risk gate. |
| Unsupported-claim rate | How often a rationale adds a number, cause, or market fact absent from the evidence snapshot. | Reject the proposal and record the failure; do not let eloquence hide missing provenance. |
FLAT compliance | Whether the model declines when data is stale, missing, ambiguous, or outside the approved scope. | Pass every forced-abstention test before shadow testing. |
| p50 / p95 / p99 latency | Time to first token, total generation time, and complete event-to-acknowledgement performance. | The worst accepted percentile must fit the strategy's event deadline with margin. |
| Peak memory and swap | Whether context growth or parallel services destabilize the Mac during release bursts. | No sustained swap or memory-pressure failure during a realistic end-to-end soak test. |
| Cost per accepted proposal | Hosted token spend or local amortization divided by proposals that survive validation. | Compare only models that pass the same evidence, safety, and latency thresholds. |
Failure-injection test matrix
| Test | Required result |
|---|---|
| Duplicate or reordered release event | At most one order candidate; event ledger remains consistent. |
| Missing macro value, wrong unit, or mixed forecast category | No trade; evidence error is recorded. |
| Model timeout, crash, or malformed JSON | No trade; risk and reconciliation services continue. |
| Stale quote, widened spread, broker disconnect, or market closed | No new order; open positions still reconcile. |
| Unknown broker acknowledgement | Search by client ID before any retry. |
| Power loss and restart | Reload current account and event state, dedupe, then resume in safe mode. |
Advance to the next testing stage only when the measured evidence clears predetermined criteria: p95 and p99 latency, data freshness, schema-rejection rate, duplicate-order count, fill slippage, maximum drawdown, risk-gate interventions, and reconciliation completeness. A backtest return alone is not an acceptance test.
What you built
The finished design uses a pinned open-weight model on a Mac mini as a private inference and control host, FXMacroData as the live macro evidence layer, and a broker as the source of executable prices and order state. You now have a model-and-memory shortlist, five reproducible download paths, and a dated cost model for deciding between local, hosted, or hybrid inference. Whichever route wins the replay, the model can propose a direction but cannot size a position, bypass stale-data checks, hold a broker token, or decide that a failed dependency is safe. That division is what turns an offline-model experiment into a testable real-time trading system.
Risk warning
Leveraged forex and CFD trading can produce rapid losses, including losses beyond an intended trade budget where account terms permit. This guide is technical education, not investment advice or a claim of profitability. Confirm broker regulation, product terms, tax treatment, and suitability in your jurisdiction, and use a practice account first.
Continue building
- Introducing FXMacroData real-time SSE streaming
- How to consume FXMacroData SSE streams
- OANDA v20 API with FXMacroData
- Hugging Face Transformers with FXMacroData
- Qwen with FXMacroData
- Mistral with FXMacroData
- DeepSeek with FXMacroData
- Hermes vs Claude vs Gemini for FX reasoning
- Point-in-time macro backtesting
- FXMacroData predictions and provenance
- FXMacroData open-source integrations
- FXMacroData AI Trading Hub
Sources and references
Mac mini and local runtimes
- Apple Mac mini technical specifications
- Apple Newsroom: M6 and M5 Pro Mac mini announcement and availability
- Apple Store M6 24GB Mac mini configuration
- Apple Mac mini Product Environmental Report
- Apple Machine Learning Research: exploring LLMs with MLX on M5
- Apple MLX-LM repository and documentation
- MLX-LM server guidance
- llama.cpp build documentation for macOS and Metal
- Ollama macOS documentation and Ollama local-only settings
- LM Studio model-download CLI
- Hugging Face CLI download and revision documentation
Open-model cards and licences
- Open Source Initiative: Open Source AI Definition
- Qwen3 4B Instruct, Qwen3.5 9B, and Qwen3.6 27B model cards
- Google Gemma 4 model documentation
- Meta Llama 3.2 3B Instruct model card and licence
- Mistral AI Ministral 3 8B official GGUF model card
- DeepSeek-R1 Distill Qwen 7B model card
- OpenAI gpt-oss introduction and model facts
- Meta publisher model-download tooling
Hosted model and consumer-plan pricing
- ChatGPT consumer plans and separate API billing
- OpenAI API models and token pricing
- Anthropic Claude API pricing
- Claude consumer plans and subscription/API billing separation
- DeepSeek API pricing and thinking-mode guidance