Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 

README.md

skycap

Trajectory capture for RL rollouts. A harness points its unchanged OpenAI client at a per-trajectory URL. skycap records every model call into a context graph — one node per message, where resamples, subagents, compaction and harness edits are forks — and, when the trajectory finishes, returns one training sample per root-to-leaf path.

skycap is its own package inside this repository and does not depend on skyrl.

Run a server

Text mode forwards to any OpenAI-compatible server and records what it sees:

cd skycap && uv sync
uv run skycap serve --upstream-url http://engine:8000/v1 --record-dir ./record

Token mode renders the prompt itself through renderers and calls a token-in/token-out engine, so the stored tokens are the ones inference saw, with logprobs, routed experts and sampling masks:

uv sync --extra tokens
uv run skycap serve --mode tokens --upstream-url http://engine:8000 \
  --tokenizer Qwen/Qwen3-8B --max-model-len 32768 \
  --sampling-overrides '{"top_k": 50}' --sampling-mask --record-dir ./record

The engine is vLLM, over its own /inference/v1/generate. Another engine's wire is a subclass of skycap.tokens.engine.VLLMEngine.

By default a reply is parsed: a thinking model's reasoning comes back as reasoning_content, and tool calls as tool_calls. Add --use-raw-content when the harness was written against a vLLM server with no reasoning or tool parser. Replies then match that server's: the completion's own text as content, with thinking inline and tool calls unparsed, and reasoning_content: null. A harness that replays content and drops reasoning_content (Terminus-2 through LiteLLM, for example) then sends each turn back unchanged, and a thinking model's history stays one path. With parsed replies, every replayed turn would lose its thinking and fork the graph.

Embed a server

A trainer can run a server in its own process instead, from the same options skycap serve takes. It gets a thread and event loop of its own:

from skycap import CaptureService

service = CaptureService(
    "http://engine:8000", mode="tokens", tokenizer="Qwen/Qwen3-8B",
    max_model_len=32768, record_dir="./record",
)
url = service.start()        # hand this to a CapturePool
...
service.stop()               # writes the trajectories still in memory

How a call reaches the model is built inside from those options. An engine with another wire passes engine= (a skycap.tokens.engine.VLLMEngine subclass), which is the one piece an embedder supplies.

Capture a rollout

from skycap import CapturePool

pool = CapturePool(["http://capture-0:8080", "http://capture-1:8080"])
async with pool.trajectory({"task": "t1", "step": 3}) as trajectory:
    run_harness(base_url=trajectory.base_url)        # any OpenAI client
    result = await trajectory.finish({"reward": 1.0})

result.status          # "finished", or "failed" if a turn couldn't be attributed exactly
for sample in result.samples:
    sample.input_ids, sample.loss_mask, sample.logprobs
    sample.routed_experts, sample.sampling_mask

Creates go round-robin over the servers, and each trajectory's URL names its server, so no router or load balancer is involved. An SDK retry (x-stainless-retry-count) gets the original call's reply rather than a second sample.

The record

Each trajectory is written once, when it ends (finish, idle TTL or graceful shutdown): a document, plus sidecars for tokens (with the text they decode to and each token's offset in it), routed experts and sampling masks. The format is specified in docs/format.md, which is what any reader, such as the viewer, implements.

Develop

cd skycap
uv sync --extra tokens
uv run pytest

Formatting and lint are the repository's (bash format.sh from the root).