Help & Documentation

Everything you need to get the most out of GetWrangler.

GetWrangler is the AI Control Tower for Spend, Routing, and Governance. For AI coding tools, it helps teams attribute usage, manage chargeback, and understand who is consuming AI resources across the organization.

Coding-tool tracking is the visibility layer. Production AI spend governance is where GetWrangler scales.

Section 1

Getting Started

What GetWrangler does

GetWrangler sits between your application and your LLM provider. It intercepts every API call, routes it to the least-cost handler, and logs verified savings — without changing your code.

Your proxy endpoint

Point your application at the GetWrangler proxy instead of your LLM provider's API directly.

https://api.getwrangler.ai

Your dashboard

Monitor savings, review routing decisions, and export your data at any time.

https://app.getwrangler.ai

Integration guide

Step-by-step setup instructions, code examples, and configuration reference.

View Integration Guide

Evaluation Period

GetWrangler offers a 14-day shadow mode evaluation. During this period, GetWrangler observes your actual traffic, establishes your savings baseline, and shows you exactly what optimization is achievable — before you activate live routing or pay anything.

Option A — Trial Mode (no Abacus account required)

Your account starts immediately using GetWrangler's operator credentials with free model routing (Gemini Flash). This lets you see routing behavior and savings projections right away.

  • Limitation — Responses use Gemini Flash regardless of your requested model
  • Limitation — Rate limiting may apply on high-volume workloads
  • Limitation — Trial is capped at 10,000 requests
  • Limitation — Savings baseline reflects free model costs, not your production model

* GetWrangler 1.x provides access to major frontier models (Claude, GPT-4o, Gemini) via a client-owned Abacus.ai account. GetWrangler 2.x will add direct native endpoint support for all major providers.

To upgrade to full evaluation at any time:

  1. Log into your dashboard at app.getwrangler.ai, navigate to the Tokens section, and click Create New Token
  2. Enter your Abacus API key — it is stored directly in our encrypted vault and never visible to GetWrangler staff
  3. Your account upgrades instantly — no operator involvement needed

Option B — Full Evaluation (recommended)

Open a free account at abacus.ai and provide your API key during onboarding OR add it yourself at any time via your dashboard. GetWrangler stores it directly in an encrypted vault.

  • Benefit — Real-world response quality matching your production model
  • Benefit — No rate limiting beyond your own Abacus plan limits
  • Benefit — Accurate savings baseline using your actual model costs
  • Benefit — Seamless transition to live routing — no credential changes needed
Your API key is stored in an encrypted vault (WorkOS). It is never logged, never visible to GetWrangler staff, and never transmitted except to your LLM provider.

At the end of your evaluation, visit your dashboard to select a plan and activate live routing. View plan options →

Section 2

Understanding Shadow Mode

What it is

When your account is in shadow mode, all requests pass through to your LLM provider unchanged. GetWrangler observes your traffic, builds your baseline, and logs what it would have saved — but takes no action.

Why it starts here

Shadow mode lets you verify GetWrangler's behavior before any routing decisions affect your application. Your 14-day evaluation window runs in shadow mode.

When it changes

Add a payment method in your dashboard and self-activate live routing when you are ready. You can see your current mode in the Routing Mode card on your dashboard.

Section 3

Reading Your Dashboard

Savings Summary

  • Total requests processed this month — count of all API calls routed through GetWrangler
  • Total verified savings — actual cost vs. baseline cost, in dollars
  • Billing projection — estimated month-end total based on current trajectory

Routing Breakdown

  • DETERMINISTIC Matched a rule or template — zero LLM cost
  • CACHE Served from cache — zero LLM cost
  • LLM Routed to least-cost eligible model
  • PASSTHROUGH Passed through unchanged (shadow mode or no cheaper option found)

Usage by API Token

  • Cost and request count — broken down by API credential
  • Separate tokens per team or service — appear as separate rows for granular attribution
  • ⚠ Spike badge — current-day spend exceeds 3× the 7-day trailing average; review for unexpected automation or runaway processes
  • Unlabeled tokens — appear as "Unlabeled / Legacy"

Vendor / Model Breakdown

Shows which AI models GetWrangler routed traffic to. Use this to understand your model mix and associated costs across providers.

Routing Mode Status

  • Shadow vs. Active — current routing mode indicator
  • Source breakdown — routing driven by global default vs. per-request header override

Invoice History

  • Monthly billing — platform fee, savings delivered, savings share, total due
  • Status — Draft → Reviewed → Sent

Section 4

Spend Monitoring

Spend Velocity

The SpendVelocity card shows your real-time burn rate at a glance. It updates continuously and projects your trajectory forward so you can act before problems become expensive.

  • Current spend rate — rolling daily average compared to your plan period
  • Projected monthly total — extrapolated from current velocity
  • Burn rate trend — accelerating, steady, or decelerating
  • Estimated overage date — the date your projected spend would reach your plan threshold at current velocity

Budget Threshold Alerts

GetWrangler sends email alerts as your spend approaches your plan limit. Alerts are sent once per day to avoid noise.

50% Informational — your spend is on track; no action required
80% Heads-up — consider whether your current plan tier is the right fit
100% Upgrade recommended — you have reached your plan threshold; upgrading improves your unit economics
Never-halt principle: GetWrangler never pauses or blocks your requests regardless of spend level. Your application keeps running. Alerts are informational — the decision to upgrade is always yours.

Anomaly Detection

The ⚠ Spike badge appears on any API token whose current-day spend exceeds 3× the 7-day trailing average. This flags runaway automations and unexpected usage before they compound.

  • What triggers it — single-day spend > 3× the trailing 7-day average for that token
  • What to do — review recent requests for that token; check for loops, retries, or new automations
  • How it clears — automatically on the next daily calculation once the rate normalizes

Section 5

Per-User Governance

Contact Management

Map your API tokens to named contacts so every dollar of AI spend is attributed to a specific person, team, or service — not just an anonymous key.

  • Token-to-person mapping — label each API token with a name, email, and department
  • Department grouping — roll up spend by team or business unit for budget reporting
  • Named contacts — alert recipients are tied to the token, not just to the admin

Per-Token Spend Thresholds

Set a daily spend threshold for each API token. When a token's daily spend crosses its threshold, GetWrangler sends an alert email directly to the associated contact.

  • Threshold level — set per token in dollars per day
  • Alert delivery — email to the named contact for that token
  • Daily deduplication — one alert per token per day, regardless of how far over threshold
  • No service interruption — thresholds are alert triggers only; requests continue unaffected

Audit & Reporting

Every API call is logged with its token, contact attribution, routing decision, and cost. Export your full spend ledger at any time as CSV for internal reporting or cost allocation.

See Section 7: Exporting Your Data for export details.

Section 6

Plans & Pricing

Plan tiers

GetWrangler offers four tiers — Starter, Growth, Pro, and Enterprise — sized by the volume of AI spend you want to put under governance. Every tier includes spend velocity monitoring, per-user tracking, department budgets, threshold alerts, anomaly detection, and the savings ledger.

  • Starter — $199/mo platform fee · up to 25 users · up to $2,500/mo AI spend governed
  • Growth — $499/mo platform fee · up to 100 users · up to $10,000/mo AI spend governed
  • Pro — $999/mo platform fee · up to 500 users · up to $50,000/mo AI spend governed
  • Enterprise — custom pricing · unlimited users · unlimited spend · dedicated support

ACH vs. Credit Card

GetWrangler charges a savings share — a percentage of verified savings — in addition to the monthly platform fee. Paying by ACH bank transfer gives you a lower savings share rate, meaning you keep more of every dollar saved.

  • ACH savings share — 20% (Starter) · 15% (Growth) · 10% (Pro/Enterprise)
  • Credit card savings share — 23% (Starter) · 18% (Growth) · 13% (Pro/Enterprise)
  • Platform fee — identical for both payment methods

Either way, you keep the substantial majority of every verified dollar saved.

Full plan comparison

See the complete feature-by-feature comparison, FAQ, and trial information.

View Pricing Page →

Section 7

Exporting Your Data

The Export Data button downloads your complete savings ledger as a CSV file.

  • Columns included — timestamp, handler type, baseline cost, actual cost, verified saving, API token
  • Row limit — capped at 10,000 rows; the file header indicates if the export is partial
  • Use cases — internal reporting, cost allocation, or audit purposes

Section 8

API Token Management

  • Multiple tokens supported — use one token per account, or separate tokens per team, service, or automation pipeline for per-token cost visibility
  • Token labels — set via your dashboard; unlabeled tokens appear as "Unlabeled / Legacy"
  • Token rotation — contact support@getwrangler.ai to rotate a token; the old token is deactivated immediately upon request
  • Per-token thresholds — set daily spend limits and alert contacts per token (see Section 5)

Per-Token Model Preferences

Each token can optionally be set to prefer a specific model for LLM-tier requests.

Three settings are available:

  • GetWrangler chooses (default) — maximum cost savings; GetWrangler selects the least-cost eligible model. This is the recommended default for most teams.
  • Prefer [model] (soft) — GetWrangler uses your preferred model when the task warrants it, and routes to a lower-cost model when it doesn't. Partial savings preserved.
  • Always use [model] (hard) — your preferred model is always used for LLM-tier requests. LLM-tier savings do not apply to this token.

To set a preference: click the arrow (▸) next to any active token in your dashboard Tokens section to expand the model preference selector.

Never-halt principle: Deterministic and cache handler savings apply regardless of model preference setting — GetWrangler always optimizes what it can.

Section 9

Getting Help

Section 10

Advanced Configuration

Native Endpoint Support

For teams with fine-tuned models or existing direct endpoint integrations, GetWrangler supports native Anthropic, OpenAI, and Google endpoints directly — so organizations with custom model investments can still benefit from GetWrangler's cost optimization, spend governance, and audit trail without abandoning their existing provider setup. Register up to 5 custom endpoints per account, each with a vault-stored API key and your own contracted pricing rates. Click Save and Test to verify the connection — verified endpoints show a 🔓 icon, unverified or failed ones show 🔒.

Where Not to Use GetWrangler

GetWrangler is built for production application traffic. A few scenarios are better handled without it.

There are two distinct use cases for AI developer tools — and GetWrangler's role differs for each.

If your goal is cost optimization: GetWrangler is not the right fit for coding assistants in general. Developers need consistent, full-capability model access without substitution or added latency. Do not route developer tool traffic through standard optimization mode.

If your goal is spend attribution, team visibility, or client chargeback: this works today for GitHub Copilot CLI — full support, including agentic/tool-use traffic, not just chat — for Cursor's chat/plan panel only (Composer/Agent mode, inline edit, and autocomplete cannot be routed through GetWrangler under any configuration), and for Claude Code and Codex CLI, which connect directly to your own Anthropic or OpenAI account for tracking and billing only, with no tier-routing on that traffic. See the Coding Tools Guide for the full breakdown. Use bypass mode on each developer token for supported tools — full model capability is preserved with zero substitution, while every routed API call is logged to a team, project, or client.

Common chargeback scenarios:

  • Track which developer or team is consuming how much
  • Allocate AI costs back to client projects or departments
  • Identify runaway usage before it hits the invoice
  • Produce per-project AI spend reports for billing

Set bypass mode on each token to preserve full model capability. Savings will show $0 — that is correct and expected. The value is attribution, not optimization.

If your total LLM API spend is under $200 per month and your goal is cost optimization, the platform fee likely exceeds what routing savings would deliver — GetWrangler becomes worth it once routing decisions start compounding across meaningful volume.

But if you're using Claude Code, GitHub Copilot, or Codex CLI to build for clients and need to invoice them for that AI usage, spend level doesn't matter — GetWrangler gives your team per-project, per-client billing attribution for one flat price, regardless of how much or little you're spending. This is true even if your total AI spend is well under $200/month.

If you're only interested in cost optimization and are below the spend threshold, you're welcome to run a free trial to see your actual numbers, but the math may not favor a paid plan yet for that use case alone.

Spend alone isn't the only factor. A client spending $500 a month on a handful of large batch requests has very little for the classifier to optimize against. Savings come from repetition and cacheable patterns. If your traffic is a small number of large, unique requests rather than many repeated ones, GetWrangler may find limited savings even at higher spend levels. Run a trial to see your actual numbers before committing to a plan.

Latency Expectations

What to expect when GetWrangler sits in your request path.

GetWrangler adds a brief classification step before your request reaches the model provider — this is what determines the most cost-effective model for the task. On average, this adds under one second of overhead, and it applies the same way whether you're using streaming or non-streaming requests. For deterministic and cached responses, you actually see lower total latency than calling your model directly, since the request never makes a round trip to the model at all. Your dashboard shows your actual average proxy overhead, upstream latency, and total latency for your account, based on your real traffic. For latency-critical workloads, bypass mode skips classification entirely and routes your request directly to your specified model for the lowest possible latency — though you trade away the cost-optimization savings for that traffic.

Cost Governance & Evaluation

Common questions about how GetWrangler handles cost control, evaluation, and integration.

Yes — GetWrangler is built on a "never halt" principle: it never pauses or blocks live traffic regardless of spend level. Budget alerts are informational only, so a cost spike triggers a notification, not a forced passthrough, throttle, or upgrade block. This is a deliberate design choice, since blocking production traffic to control costs creates a worse outage risk than the cost overrun itself.

Yes — GetWrangler supports shadow mode, a 14-day evaluation period where all requests pass through to your LLM provider completely unchanged. GetWrangler observes your traffic in the background, classifies it, and projects what optimization would have saved — without altering what your application actually receives. You see projected savings and routing behavior before you activate live optimization, at no risk to production traffic.

GetWrangler provides AI spend governance with configurable thresholds and real-time alerts, without ever forcing a pause in service (see "never halt" above). Governance features are available starting on the Pro tier ("AI Spend Governance"), with the Enterprise tier ("AI Control Plane") adding organization-wide policy and routing control across teams.

Yes — GetWrangler works as a transparent proxy in the API call path, so it adds cost tracking, classification, and routing without requiring engineers to rewrite application code or change how they call the model API. Teams get spend visibility and optimization without adding integration work to their existing workflow.

Yes. GetWrangler supports bypass mode, a per-token setting available anytime — including during your 14-day trial, with no subscription required — that routes traffic through full tracking and attribution infrastructure while preserving complete model capability with zero substitution. This is commonly used with GitHub Copilot CLI and Cursor's chat/plan traffic, where the value is spend visibility and attribution rather than automatic cost optimization. Claude Code and Codex CLI work differently: they connect through dedicated direct-passthrough routes that forward requests unchanged to your own provider credentials, with tracking-only behavior built in by design rather than a toggle you set.

Yes — set your token's Model Preference to Let each call decide in your dashboard, then control routing per request via the model field. Two request shapes are supported:

  • Explicit model — name an exact model_id (e.g. "claude-haiku-4-5-20251001") and GetWrangler executes exactly that model for the call, no substitution.
  • Anchored auto-route — prefix the model with getwrangler-auto: (e.g. "getwrangler-auto:claude-haiku-4-5-20251001") and GetWrangler picks the best fit for the task, but never anything pricier than the named model. Savings are measured against that model's price.

A bare "getwrangler-auto" with no anchor model is not valid — auto-routing always requires a cost ceiling — and returns a 400 error.