TokenPolice
Docs
Core concepts

Identity: user_id, paid_plan, session_id

The tags that let a rule target a slice of your traffic — what each one means, what happens if you skip it, and how to set them.

Rules don't act on your whole app at once — they target slices of traffic. The tags below define those slices. Attach them and a rule can say "cap free-plan users", "stop this one runaway conversation", or "watch what the summarizer step costs".

None of them are required. Skip them all and TokenPolice still meters everything; you just can't tell one user, plan, or conversation from another.

The fields

FieldWhat it isIf you don't set it
user_idWhich end user made the call"anonymous" — all spend lands in one pool, so only app-wide limits work
paid_planWhich pricing tier that user is on"free"
session_idWhich conversation or run this call belongs toa fresh id per run, so nothing threads together
workflow nameWhat this run is"support_agent", "nightly_report""default_workflow"
span nameWhat this one LLM step is — "summarize", "route"the model's name
metadataAny other tags you want to slice by — tenant_id, feature, regionempty

user_id and paid_plan are the two that make per-user and per-tier budgets possible. session_id is what turns a multi-turn chat into one costed conversation. The names are what make the dashboard readable — and all of them are things a rule can match on.

paid_plan is compared exactly. It isn't trimmed and isn't lowercased, so "Free", " free " and "free" are three different plans: a rule looking for free won't match the first two, and a rule that pools by plan gives each spelling its own budget. Normalize once, in your app — str(tier or "free").strip().lower() / String(tier ?? "free").trim().toLowerCase() — and use that same spelling when you write the rule.

The stable-session_id contract

session_id is the one field you have to be deliberate about: pass the same id on every turn of the same conversation. A new id per turn isn't a smaller mistake — it means the conversation never exists as far as TokenPolice is concerned, and per-conversation caps can never fire.

Use whatever your app already has: a chat id, a thread id, a support ticket number. What the SDK does with it:

  • Trims it and strips control characters.
  • Caps it at 200 characters.
  • If it's empty or missing, falls back to a fresh id for that run alone.

If your app really is one-shot per request, leaving it out is a fair choice — just know that conversation grouping and per-conversation limits won't apply.

Setting them

The usual way is to wrap your entrypoint and let it read the values off the call.

import token_police as tp

@tp.workflow(name="support_agent")
def handle(user_id: str, paid_plan: str, session_id: str, query: str):
    ...   # every LLM call inside carries these three
import * as tp from "token-police";

const handle = tp.workflow({ name: "support_agent" }, async (req: {
  userId: string; paidPlan: string; sessionId: string; query: string;
}) => {
  // every LLM call inside carries these three
});

Python reads arguments by name — a parameter literally called user_id, paid_plan, session_id or metadata. Node reads the same keys off the first argument when it's an object (userId, paidPlan, sessionId; snake_case works too). Either way you can also pass the values statically, or open a session(...) block instead of wrapping a function — the full API is in Wrapping your app.

To name one step inside a run, call set_span_name just before it.

Nested scopes inherit the enclosing scope's ids and merge its metadata, so you set identity once at the top and everything below is tagged.

Next