loopl¶
native swift · edge inference
Agents that run on the phone.
No cloud, no login.¶
loopl is a Swift agent SDK built for on-device inference — the Strands agent loop, rewritten for Metal and the Neural Engine. Pick a model from Hugging Face, hand the agent a few tools, and let it run until the job is done. The same loop powers the free loopl iOS app.
Coming from Strands Agents? See how loopl compares →
import Loopl
import LooplRuntimes
let model = MLXModel() // Metal, on the phone
try await model.load(spec, from: modelDir) { _ in } // mlx-community/Qwen3-1.7B-4bit, 968 MB on disk
let agent = Agent(model: model,
tools: [CurrentTimeTool(), CalculatorTool()],
systemPrompt: "You are a concise assistant.")
for try await event in agent.stream("What time is it in Istanbul, and what's 23*47?") {
switch event {
case .text(let t): print(t, terminator: "")
case .toolStarted(let call): print("→ \(call.name)")
case .toolFinished(let r): print("← \(r.content)")
case .metrics(let m) where m.finished: print("
\(Int(m.tokensPerSecond)) tok/s")
default: break
}
}
Swift SDK¶
Agent, Tool, Model, Message, AgentEvent — a small, typed surface that mirrors Strands
naming where it is natural, and feels like Swift everywhere else. Actors, async streams,
Sendable all the way down.
On-device runtimes¶
MLX Swift for 4-bit Hugging Face models on the GPU, Apple Foundation Models on iOS 26 with
zero download, and GGUF / llama.cpp on the roadmap. One Model protocol, swap at runtime.
Unlimited agent loop¶
No max_iterations = 8. The loop runs until the model stops calling tools or you cancel.
Optional Budget caps for iterations, tool calls, wall time and tokens, plus a LoopDetector that
warns on identical repeated calls — and never silently stops.
Why on-device?¶
What you get¶
- Privacy by construction — prompts, tool results and notes never leave the phone.
- Latency you can feel — first token in well under a second on an A18 Pro with a 1.7B model.
- Works offline — a model already on disk answers in airplane mode.
- Free to run — no API key, no metered tokens, no vendor bill.
What to expect¶
- Models are 0.6B–4B parameters at 4-bit; 7B+ is a warning, not a default.
- Context is RAM-bound: 8k tokens is the sweet spot on an 8 GB iPhone.
- Tool calling quality varies by family — Qwen3 and Llama 3.2 are the reliable ones.
- See the models table for size, context and tool-call support per model.
One struct to a tool¶
struct Clock: Tool {
let name = "current_time"
let description = "Current date & time, optionally for an IANA time zone"
let parameters = ToolSchema.object(["timezone": ToolSchema.string("e.g. Europe/Istanbul")])
func call(_ args: [String: JSONValue]) async throws -> String {
let tz = TimeZone(identifier: args["timezone"]?.stringValue ?? "") ?? .current
return Date().formatted(.dateTime.timeZone(tz))
}
}
The schema is handed to the model's own chat template (Qwen/Hermes <tool_call>, Llama 3.x) or listed in the
instructions for Apple Foundation Models and Gemma-class models; ToolCallParser normalises all three dialects —
you write the tool once.
Tools & JSON schema → Browse the models → Read the privacy policy →