ToolPiper API
ToolPiper is the local AI server that powers ModelPiper. It runs on your Mac, manages inference engines, and exposes an OpenAI-compatible API at localhost:9998.
http://localhost:9998/v1SpecGET /v1/openapi.jsonHow It Works
ModelPiper is the web app you're looking at right now. It talks to ToolPiper, a native macOS app running in the background. ToolPiper manages the actual AI engines — llama.cpp for LLMs, FluidAudio for speech, CoreML for images, and more.
When you create a Provider in ModelPiper, you're setting up a configuration that pairs an AI model with a specific engine. For example: Use the Llama 3.2 3B model via llama.cpp or Use Parakeet for speech-to-text via FluidAudio. Each provider becomes a usable endpoint on the ToolPiper API.
ModelPiper uses an internal session key to stay connected to ToolPiper — you don't need to think about that. But if you want to build your own app on top of ToolPiper, you'll need a Developer Token.
Developer Tokens
A dev token lets you use ToolPiper from your own code, just like an OpenAI API key. Create one, drop it into any OpenAI-compatible SDK, and point the base URL at localhost:9998/v1. That's it.
Claude Code Zero-config
ToolPiper ships a bundled claude-tp helper that bridges Claude Code to any provider you've configured here. Ask ToolPiper in chat to install it, or call the claude_code_install MCP tool: it symlinks claude-tp onto your PATH, mints a scoped inference token, and writes the helper's own config (never ~/.claude/). Run claude-tp in place of claude and the endpoints below are available. Most of them need no plan; the badge on a row says which one the rest need.
New to all of this? See the comparison vs Ollama / vLLM / LM Studio.
claude_code_install MCP tool. It's idempotent and self-healing: re-running refreshes the symlink, token, and config. Remove it with claude_code_uninstall. Installing the helper needs /v1/tokens stays free, and so does pointing any Anthropic-compatible client at the proxy yourself. Anthropic Proxy Backend Phase 1 plumbing
ToolPiper will expose POST /v1/messages as an Anthropic-shape proxy so Claude Code (and any Anthropic-compatible client) can use any provider you've configured. Pick the global backend below, or bind a specific provider per token in the table above.
Quick Start
Drop-in replacement for the OpenAI SDK — just change the base URL and API key. Questions? @ModelPiper on X.
MCP Server 422 tools
ToolPiper is also an MCP server. Install categories individually to control which tools your AI client sees — saves context tokens.
See MCP docs for install commands, profiles, and full tool reference.
Endpoints
Inference
OpenAI-compatible inference endpoints
- GET
/v1/modelsList available models - POST
/v1/chat/completionsCreate chat completion - POST
/v1/embeddingsCreate embedding (OpenAI-compatible) - POST
/v1/benchmark/llmBenchmark on-disk text-generation models
Audio
Speech-to-text and text-to-speech
- POST
/v1/audio/transcriptionsTranscribe audio - POST
/v1/audio/speechText-to-speech - POST
/v1/audio/recordRecord an audio clip
Image
Image processing (upscale)
- POST
/v1/benchmark/upscaleRun PiperSR benchmark suite
RAG
Retrieval-Augmented Generation — collections, ingestion, and semantic search
Cloud Proxy
Keychain-backed cloud API proxy
Models
Model management, downloads, and HuggingFace integration
- GET
/v1/models/installedList downloaded models - GET
/v1/models/{modelId}Get model details - DELETE
/v1/models/{modelId}Delete a model - GET
/v1/models/storageDisk usage for models - GET
/v1/models/searchSearch HuggingFace for models - POST
/v1/models/downloadDownload a model from HuggingFace - GET
/v1/models/infoInspect a model candidate from any source - POST
/v1/models/acquireAcquire a model component from any source - GET
/v1/models/downloadsList active downloads - DELETE
/v1/models/downloads/{downloadId}Cancel a download - POST
/v1/models/scanScan for new models - GET
/v1/models/hf/{owner}/{repo}/filesList files in a HuggingFace repo
Model Configs
Curated model presets with availability status
- GET
/v1/model-configsList model presets - POST
/v1/model-configs/installInstall a model by preset ID
Engine
Inference engine control and model state
- GET
/v1/engine/statusEngine and backend status - POST
/v1/engine/loadLoad a model into the engine - POST
/v1/engine/unloadUnload a model or stop the engine - GET
/v1/models/statePer-model runtime states - POST
/v1/models/reloadReload a llama-server-backed model
Tokens
Developer token management (all tiers; requires the tokensManage scope)
- GET
/v1/tokensList developer tokens - POST
/v1/tokensCreate a developer token (all tiers; requires the tokensManage scope) - PATCH
/v1/tokens/{tokenId}Update a developer token's metadata - DELETE
/v1/tokens/{tokenId}Revoke a developer token
Apple
Apple-native framework tools (Vision, NLP) — no model downloads required
- POST
/v1/apple/ocrRecognize text in an image - GET
/v1/apple/ocr/languagesList supported OCR languages - POST
/v1/apple/barcodeDetect barcodes and QR codes - POST
/v1/apple/classifyClassify image content - POST
/v1/apple/face-detectDetect faces - POST
/v1/apple/saliencyDetect salient regions - POST
/v1/apple/rectanglesDetect rectangles - POST
/v1/apple/feature-printGenerate image feature vector - POST
/v1/apple/body-poseDetect human body poses - POST
/v1/apple/hand-poseDetect hand poses - POST
/v1/apple/animalsDetect animals - POST
/v1/apple/horizonDetect horizon angle - POST
/v1/apple/documentDetect document boundaries - POST
/v1/apple/nlp/languageDetect language - POST
/v1/apple/nlp/sentimentAnalyze sentiment - POST
/v1/apple/nlp/entitiesNamed entity recognition - POST
/v1/apple/nlp/tokenizeTokenize text - POST
/v1/apple/nlp/lemmatizeLemmatize text - POST
/v1/apple/nlp/posPart-of-speech tagging
System
Health checks, resource monitoring, events, and licensing
- GET
/statusHealth check - GET
/session-tokenObtain ambient bearer token - GET
/v1/permissionsmacOS permission snapshot - GET
/v1/system/resourcesGPU, RAM, and ANE utilization - GET
/v1/eventsSSE event stream - GET
/v1/licenseSubscription tier and features - GET
/v1/openapi.jsonOpenAPI specification - GET
/v1/audit/toolsTool dispatch audit log - DELETE
/v1/audit/toolsPurge the audit trail - GET
/v1/audit/tools/rollupPermanent audit rollup - GET
/v1/snippets/match/statusSnippet match diagnostics
Configurations
Endpoint configuration management
- GET
/v1/configurationsList endpoint configurations - POST
/v1/configurationsCreate an endpoint configuration - PUT
/v1/configurations/{configId}Update an endpoint configuration - DELETE
/v1/configurations/{configId}Delete an endpoint configuration
Logs
Log ingestion, querying, and real-time streaming
- GET
/v1/logsQuery log entries - POST
/v1/logsIngest log entries - GET
/v1/logs/streamReal-time log stream (SSE) - POST
/v1/logs/clearClear all log entries - POST
/v1/logs/exportExport logs to file
Templates
Workflow templates
- GET
/v1/workflow-templatesList workflow templates
Browser
CDP-based browser automation
- GET
/v1/browser/statusBrowser connection status - POST
/v1/browser/connectConnect to a browser via CDP - POST
/v1/browser/disconnectDisconnect from browser - GET
/v1/browser/pagesList open browser pages - POST
/v1/browser/select-pageSelect a browser page - GET
/v1/browser/snapshotAccessibility tree snapshot - GET
/v1/browser/screenshotPage screenshot - GET
/v1/browser/consoleConsole messages - POST
/v1/browser/record/startStart recording user actions - POST
/v1/browser/record/stopStop recording user actions - GET
/v1/browser/record/streamRecording event stream (SSE) - POST
/v1/browser/network/enableEnable network logging - POST
/v1/browser/network/disableDisable network logging - GET
/v1/browser/networkList captured network entries - DELETE
/v1/browser/networkClear captured network entries - GET
/v1/browser/network/statusNetwork logging status - GET
/v1/browser/network/{requestId}/bodyGet response body for a captured request - POST
/v1/browser/trace/startStart performance trace - POST
/v1/browser/trace/stopStop performance trace - GET
/v1/browser/metricsGet current performance metrics - GET
/v1/browser/trace/statusTrace status - POST
/v1/browser/intercept/enableEnable network interception - POST
/v1/browser/intercept/disableDisable network interception - GET
/v1/browser/intercept/statusInterception status - GET
/v1/browser/mocksList all mock rules - POST
/v1/browser/mocksCreate a mock rule - DELETE
/v1/browser/mocksDelete all mock rules - PUT
/v1/browser/mocks/{id}Update a mock rule - DELETE
/v1/browser/mocks/{id}Delete a mock rule - POST
/v1/browser/coverage/startStart code coverage collection - POST
/v1/browser/coverage/stopStop code coverage and get report - GET
/v1/browser/coverage/statusCoverage status - GET
/v1/browser/storageRead browser storage - DELETE
/v1/browser/storageClear browser storage - POST
/v1/browser/storage/cookieSet a cookie - POST
/v1/browser/storage/localSet a localStorage item - POST
/v1/browser/storage/sessionSet a sessionStorage item - POST
/v1/browser/webauthn/enableEnable virtual authenticator - POST
/v1/browser/webauthn/disableDisable virtual authenticator - GET
/v1/browser/webauthn/statusWebAuthn status - GET
/v1/browser/webauthn/credentialsList virtual credentials - DELETE
/v1/browser/webauthn/credentials/{credentialId}Delete a virtual credential - POST
/v1/browser/webauthn/verifySet user verification state - POST
/v1/browser/autofill/credit-cardTrigger credit card autofill - POST
/v1/browser/autofill/addressTrigger address autofill - GET
/v1/browser/full-statusExtended browser status with pages and channels - POST
/v1/browser/actionPerform a browser action (click, fill, navigate, etc.) - POST
/v1/browser/assertPerform a browser assertion - POST
/v1/browser/healHeal a broken selector - POST
/v1/browser/evaluateExecute JavaScript in the browser - POST
/v1/browser/resizeResize browser viewport - POST
/v1/browser/dialogHandle a JavaScript dialog (alert, confirm, prompt) - POST
/v1/browser/close-tabClose a browser tab - POST
/v1/browser/new-tabOpen a new browser tab - GET
/v1/browser/channelsList available Chrome channels (dev, canary, stable)
Testing
PiperTest — visual test session management and execution
- POST
/v1/test-sessions/{sessionId}/exportMax $49/mo Export a test session to Playwright or Cypress code - POST
/v1/browser/probe/scanScan the current page for interactive elements (PiperProbe)
Pose
Human pose estimation via Apple Vision / CoreML, including on-demand real-time 60fps skeleton streaming via WebSocket
Stream
Real-time stream processing
- POST
/v1/stream/startStart a stream processing session - POST
/v1/stream/stopStop the active stream processing session - GET
/v1/stream/statusGet stream session status - POST
/v1/stream/outputDrain a stream session's processed results (pull delivery)
Scrape
CDP-based web page scraping with framework-aware readiness. Extracts content in up to 7 formats (markdown, text, readability, axTree, html, links, screenshot) from a single page load
Video
Video creator pipeline — settings, screenplays, recording, rendering, and narration
- POST
/v1/media/hostMint a hosted-media artifact - GET
/v1/media/clip/{id}Poll a clip export job - GET
/v1/media/record/{id}Poll a recording job - POST
/v1/media/record/{id}/stopStop a recording
Voice Chat
Persistent voice conversation with model selection, conversation memory, and sentence-level TTS streaming
- GET
/v1/voice-chat/settingsGet voice chat settings - PUT
/v1/voice-chat/settingsUpdate voice chat settings - POST
/v1/voice-chat/sessionCreate a voice chat session - DELETE
/v1/voice-chat/sessionEnd the active voice chat session
Pipeline
Workflow pipeline orchestration
- POST
/v1/pipeline/runExecute a workflow pipeline - POST
/v1/pipeline/run/cancelCancel the active pipeline run
Tool Permissions
MCP tool permission policies
- GET
/v1/tool-permissionsList MCP tool permission policies - PUT
/v1/tool-permissionsSet MCP tool permission policies
Tools
Unified tool catalog, retrieval probes, and session-scoped client-tool registration (PiperMatch / ToolGate plumbing)
- GET
/v1/tools/catalogSnapshot of the active ToolGate catalog (DEBUG only) - POST
/v1/tools/probeRank a candidate tool against PiperMatch on a set of cases - POST
/v1/tools/register-client-catalogRegister session-scoped client tools into the unified index - POST
/v1/tools/unregister-client-catalogUnregister a session's client tool catalog - GET
/v1/tools/nativeNative tool catalog (unfiltered)
Jobs
Shared async-job substrate — list, inspect, and cancel long-running verb jobs (chat-to-feature foundation Phase 2)
- GET
/v1/jobsList async-job records - GET
/v1/jobs/{jobId}Fetch a single async-job record - POST
/v1/jobs/{jobId}/cancelCancel an in-flight async job
Assets
Provenance-linked typed-Asset substrate — byte→typed-Asset rich ingress (unified-memory-media-substrate Phase 3)
- POST
/v1/assets/importImport raw bytes as a typed Asset
Projects
Media-project source pool — the SPA reads a project container and references / removes / reorders its source files (docs/architecture/media-project-file.md)
- GET
/v1/projects/{id}Fetch a media-project container - POST
/v1/projects/{id}/sourcesReference a file into a project's source pool - POST
/v1/projects/{id}/sources/reorderReorder a sibling group in a project's source pool - DELETE
/v1/projects/{id}/sources/{sourceId}Remove a source from a project's pool
Vault
Encrypted Writing Vault — decrypted note/folder CRUD over loopback while natively unlocked; every content route answers 423 vault_locked otherwise. Unlock and passphrase flows are native-only and have no routes (docs/plans/encrypted-writing-vault.md Phase 4)
- GET
/v1/vault/statusVault presence and lock state - GET
/v1/vault/treeThe whole decrypted vault tree - POST
/v1/vault/notesCreate an encrypted note - GET
/v1/vault/notes/{id}Read a decrypted note - PUT
/v1/vault/notes/{id}Overwrite a note's content - DELETE
/v1/vault/notes/{id}Delete a note - POST
/v1/vault/notes/{id}/moveMove and/or rename a note - POST
/v1/vault/foldersCreate a folder - DELETE
/v1/vault/folders/{id}Delete a folder recursively - POST
/v1/vault/folders/{id}/renameRename a folder in place - POST
/v1/vault/folders/{id}/moveMove a folder under a new parent - POST
/v1/vault/lockLock the vault - POST
/v1/vault/assetsImport an encrypted asset (exact bytes) - POST
/v1/vault/assets/import-urlImport an asset from a URL (or keep it as a live link) - GET
/v1/vault/assets/{id}Decrypt an asset for rendering
Capture
Screen and color capture utilities
- POST
/v1/screenshotTake a screenshot - POST
/v1/color/pickPick a color from the screen - POST
/v1/camera/captureCapture a webcam still
Files
File-system utilities — archive, PDF extraction, code search
- POST
/v1/archiveCreate, extract, or list an archive - POST
/v1/pdf/extractExtract text, pages, metadata, or images from a PDF - POST
/v1/filesystem/edit-fileExact-string edit of a text file - POST
/v1/filesystem/metadataNative media metadata for one file - POST
/v1/filesystem/move-fileMove or rename a file or directory - POST
/v1/search/codeGrep-style code search across a directory
Web
Outbound HTTP and web search
- POST
/v1/web/searchRun a web search - POST
/v1/http/requestPerform an outbound HTTP request
Translation
On-device text translation (Apple Translation framework)
- POST
/v1/translateTranslate text via the Apple Translation framework
Utilities
General macOS utilities — clipboard history, timers, QR codes, image transforms
- POST
/v1/clipboardRead, write, or manage clipboard history - POST
/v1/timersStart, list, cancel, or clear timers - POST
/v1/qr/generateGenerate a QR code PNG - POST
/v1/image/transformResize, crop, rotate, or convert an image - POST
/v1/image/segmentSegment an image Asset into per-region masks (non-minting)
Agent
Interactive agent kernel — start a turn on a session and consume its per-session SSE event stream (token deltas, tool-call proposals, tool results, cost frames)
- POST
/v1/agent/turnStart one interactive agent-kernel turn - POST
/v1/agent/turn/cancelStop the session's in-flight turn - GET
/v1/agent/streamPer-session agent event stream (SSE) - POST
/v1/agent/approveAnswer a parked interactive tool approval - POST
/v1/agent/client-resultReturn a delegated client tool's result - GET
/v1/agent/touchedList agent-touched file paths under a workspace root - GET
/v1/agent/touched/reviewThe durable review payload for one agent-touched file - POST
/v1/agent/touched/clearClear a reviewed agent touch
Task
The agent's per-session todo list — create/update/complete tasks the kernel tracks across a session, and read the list back
- POST
/v1/task/createCreate a task in a session's todo list - POST
/v1/task/updateUpdate a task's status - POST
/v1/task/completeMark a task completed - GET
/v1/task/listList a session's tasks
MCP
Model Context Protocol (Streamable HTTP transport)
- GET
/mcpGET not supported for MCP - POST
/mcpHandle MCP JSON-RPC request - DELETE
/mcpTerminate an MCP session - GET
/v1/mcp-serversList MCP servers - GET
/v1/mcp-servers/healthMCP server health - POST
/v1/mcp-servers/{name}/restartRestart an MCP server - POST
/v1/mcp-servers/{name}/reload-toolsReload an MCP server's tools - PUT
/v1/mcp-servers/{name}Add or update an MCP server
Anthropic
Anthropic Messages API proxy — drop-in backend for Claude Code and any Anthropic-shape client
- POST
/v1/messagesAnthropic Messages — drop-in proxy for Claude Code - POST
/v1/messages/count_tokensEstimate input tokens for an Anthropic Messages request - GET
/v1/settings/anthropic-proxyGet the global Anthropic Proxy backend setting - PATCH
/v1/settings/anthropic-proxyUpdate the global Anthropic Proxy backend setting
Claude Code
Zero-config integration with Anthropic's Claude Code CLI
- GET
/v1/claude-code/helper-statusclaude-tp helper status
SERP
Search-engine results capture, rank tracking, and autocomplete keyword expansion (async job + poll)
- POST
/v1/serp/autocompleteExpand a seed phrase via Google Suggest - POST
/v1/serp/searchStart a SERP capture job - GET
/v1/serp/search/{jobId}Poll a SERP capture job - POST
/v1/serp/rank-checkStart a SERP rank-check job - GET
/v1/serp/rank-check/{jobId}Poll a SERP rank-check job
YouTube
YouTube transcript extraction
Connected Apps
Connected OAuth/bearer clients — listing, pending-consent management, blocking, and revocation
- GET
/v1/connected-appsList connected apps - GET
/v1/connected-apps/blockedList blocked apps - GET
/v1/connected-apps/pending-consentsList pending OAuth consents - POST
/v1/connected-apps/{id}/blockBlock a connected app - POST
/v1/connected-apps/{id}/revokeRevoke a connected app's token - POST
/v1/connected-apps/blocked/{id}/unblockUnblock a blocked client
Action Piper
macOS UI automation — VLM grounding resolver, circuit breaker, and accessibility telemetry
- POST
/v1/action-piper/resolveResolve a click target to coordinates - POST
/v1/action-piper/circuit/resetReset the Apple FM circuit breaker - GET
/v1/action-piper/telemetry/summaryAccessibility outcome telemetry
Legacy: Ollama-compatible API NewLegacy dialect
ToolPiper serves the Ollama wire dialect as a legacy surface — a migration on-ramp, not a destination. Two mounts, one module: the opt-in loopback-only listener on :11434 (Settings → General, off by default, no auth — upstream's own posture), and this documented :9998 mount at /legacy/ollama/api/* behind the normal bearer middleware (set a configurable client's base URL to http://127.0.0.1:9998/legacy/ollama). Every response carries RFC 9745 "Deprecation" and a Link rel="successor-version" header — the surface was born deprecated, and each endpoint below names its /v1/ successor, which is where new integrations should land. Wire shapes are pinned to recorded fixtures from Ollama 0.23.4. Modelfile semantics (create, push, copy, blobs) are permanently rejected with guidance naming the successor route. The layer is deliberately disposable: it exists while it earns its keep, then it gets deleted.
http://127.0.0.1:11434opt-in listener — Settings → General, off by default, loopback-onlyBase URLhttp://127.0.0.1:9998/legacy/ollamaalways mounted, bearer-authenticated- GET
/legacy/ollama/api/versionOllama-dialect version probe — successor: GET /version - GET
/legacy/ollama/api/tagsOllama-dialect installed-model list — successor: GET /v1/models/installed - GET
/legacy/ollama/api/psOllama-dialect loaded-model list — successor: GET /v1/models/state - POST
/legacy/ollama/api/showOllama-dialect per-model detail — successor: GET /v1/models/installed - POST
/legacy/ollama/api/chatOllama-dialect chat (NDJSON stream) — successor: POST /v1/chat/completions - POST
/legacy/ollama/api/generateOllama-dialect generate (NDJSON stream) — successor: POST /v1/chat/completions - POST
/legacy/ollama/api/embedOllama-dialect embeddings — successor: POST /v1/embeddings - POST
/legacy/ollama/api/embeddingsOllama-dialect embeddings (alias upstream itself deprecates) — successor: POST /v1/embeddings - POST
/legacy/ollama/api/pullOllama-dialect model pull (NDJSON progress) — successor: POST /v1/models/download - DELETE
/legacy/ollama/api/deleteOllama-dialect model delete — successor: DELETE /v1/models/{id}
Credentials
- GET
/v1/credentialsList credentials - POST
/v1/credentials/api-keyCreate or update an API-key credential - DELETE
/v1/credentials/api-key/{credentialId}Delete an API-key credential - POST
/v1/credentials/oauth/connectConnect an OAuth grant - POST
/v1/credentials/oauth/connect/service-accountConnect an OAuth grant via service account - DELETE
/v1/credentials/oauth/{providerId}Disconnect an OAuth grant