Skip to main content
cycls.LLM holds everything about how an agent runs. It is immutable, so each method returns a new builder and one base config can be branched safely.

Choosing a model

Model strings are always vendor/model. Anything that is not anthropic/ goes through the Chat Completions adapter, so base_url points it at the endpoint and api_key supplies the key.
The prefix selects the reasoning dialect, so use the vendor prefix that matches the API you are calling, even when you self-host that model. See Self-hosted and open models for serving your own.
Model identifiers change as vendors ship new versions. Check the vendor’s documentation for the current name before pinning one in production.
Keys come from the environment (ANTHROPIC_API_KEY, OPENAI_API_KEY) or from .api_key(). Use .headers() for endpoints that authenticate outside the bearer token, such as a Modal proxy or Cloudflare Access.

Reasoning

.thinking() is one unified control, translated into each vendor’s dialect.
A vendor with no mapping, such as a host prefix like modal or vllm, gets no reasoning parameter and prints one warning. Use .extra_body() for those.
.extra_body() merges after the built-in mapping, so your keys win. It is the escape hatch for any parameter Cycls does not model.

Budgets and cost

.context() is what decides when compaction starts, so set it to match the model you actually run. .price() takes USD per million tokens. With prices set, every turn logs its cost and cycls cost and cycls sql can slice spend by user, chat or model. Without it, costs report as zero.

Long conversations

The loop keeps a chat inside the context window without ending it. Before each model call, and once more when a run ends, it compares the last request’s input tokens with a trigger. Past the trigger, the loop frees room in two steps:
  1. Clear. Old tool results, and tool arguments longer than 1,000 characters, are replaced with a short placeholder. Paths and commands stay, so the model still knows what it did. A loaded skill is never cleared. This makes no model call.
  2. Summarize. When clearing would free less than 25% of the window, the model summarizes everything before the recent 30% instead. The request reuses the loop’s system prompt and tools, so the provider serves most of it from cache. The user sees “Summarizing earlier messages to keep this chat going…” while it runs.
A chat that resumes after 20 minutes idle is cleared before its first call when that frees 10% of the window. The provider’s cache has mostly expired by then, so clearing costs little in cache hits. The system prompt, including AGENT.md and the skill catalog, is never cleared or summarized. The saved transcript is never rewritten either, so users always see the whole conversation. If a summary fails or takes longer than five minutes, the model stops seeing the older turns, gets a line saying they were dropped, and the chat continues. When the provider still reports that the context is full, the loop summarizes and replays the turn, once per run. One read returns at most 50,000 characters, so a single large file or attachment cannot fill the window by itself.

Vision and text-only models

Attachments are sent as base64 media by default. Text-only models reject that, so turn vision off and the file stays in the workspace with a note naming it, which the model can then open with a tool.
brave is the portable search and fetch pair. It works on any model and needs BRAVE_API_KEY. Without that key it falls back to the provider’s native search where one exists. native forces the provider’s server-side search, which today means Anthropic only.

Full builder reference

Next steps

Tools

Built-in tools, custom handlers and approvals.

Self-hosted models

Point an agent at vLLM, SGLang or a private endpoint.