GGUF / llama.cpp (v2)¶
GGUFModel exists as a stub today so the catalog and the runtime picker can describe GGUF entries:
GGUFModel.availability() // RuntimeAvailability(available: false, reason: "llama.cpp (GGUF) is planned for v2 — not in this build.")
try await GGUFModel().load(spec, from: dir) { _ in } // throws ModelError with the same reason
Why v2 and not v1¶
- llama.cpp no longer ships a
Package.swift; the SwiftPM path is the xcframework from each release (llama-bNNNNN-xcframework.zip, ~60 MB) wrapped as a localbinaryTarget. - The C API (
llama.h) has no chat-template or tool-call handling. That lives in C++common/chat.h(common_chat_templates_init,common_chat_parse), which is not exported — so either a C shim is built into the xcframework or the rendering/parsing is done in Swift. loopl'sPromptBuilder+ToolCallParserare that Swift side already, which is why the MLX path shipped first. n_ctxallocates the KV cache up front; the right value has to be derived fromos_proc_available_memory()after the model loads.
What it will give you¶
- The widest model coverage (every Q4_K_M on the Hub; ~10–15 % larger than MLX 4-bit).
- Single-file downloads that resume trivially.
- Metal on A18 Pro, same GGML backend as the Mac.
Planned catalog entries (none gated): Qwen3 0.6B/1.7B/4B, Llama-3.2 1B/3B, gemma-3-1b-it, Phi-4-mini, SmolLM2/3 — the GGUF rows appear in the models table once the runtime exists.