Trajectory capture for RL rollouts. A harness points its unchanged OpenAI client at a per-trajectory URL. skycap records every model call into a context graph — one node per message, where resamples, subagents, compaction and harness edits are forks — and, when the trajectory finishes, returns one training sample per root-to-leaf path.
skycap is its own package inside this repository and does not depend on
skyrl.
Text mode forwards to any OpenAI-compatible server and records what it sees:
cd skycap && uv sync
uv run skycap serve --upstream-url http://engine:8000/v1 --record-dir ./recordToken mode renders the prompt itself through
renderers and calls a
token-in/token-out engine, so the stored tokens are the ones inference saw,
with logprobs, routed experts and sampling masks:
uv sync --extra tokens
uv run skycap serve --mode tokens --upstream-url http://engine:8000 \
--tokenizer Qwen/Qwen3-8B --max-model-len 32768 \
--sampling-overrides '{"top_k": 50}' --sampling-mask --record-dir ./recordThe engine is vLLM, over its own /inference/v1/generate. Another engine's wire
is a subclass of skycap.tokens.engine.VLLMEngine.
By default a reply is parsed: a thinking model's reasoning comes back as
reasoning_content, and tool calls as tool_calls. Add --use-raw-content
when the harness was written against a vLLM server with no reasoning or tool
parser. Replies then match that server's: the completion's own text as
content, with thinking inline and tool calls unparsed, and
reasoning_content: null. A harness that replays content and drops
reasoning_content (Terminus-2 through LiteLLM, for example) then sends each
turn back unchanged, and a thinking model's history stays one path. With parsed
replies, every replayed turn would lose its thinking and fork the graph.
A trainer can run a server in its own process instead, from the same options
skycap serve takes. It gets a thread and event loop of its own:
from skycap import CaptureService
service = CaptureService(
"http://engine:8000", mode="tokens", tokenizer="Qwen/Qwen3-8B",
max_model_len=32768, record_dir="./record",
)
url = service.start() # hand this to a CapturePool
...
service.stop() # writes the trajectories still in memoryHow a call reaches the model is built inside from those options. An engine
with another wire passes engine= (a skycap.tokens.engine.VLLMEngine
subclass), which is the one piece an embedder supplies.
from skycap import CapturePool
pool = CapturePool(["http://capture-0:8080", "http://capture-1:8080"])
async with pool.trajectory({"task": "t1", "step": 3}) as trajectory:
run_harness(base_url=trajectory.base_url) # any OpenAI client
result = await trajectory.finish({"reward": 1.0})
result.status # "finished", or "failed" if a turn couldn't be attributed exactly
for sample in result.samples:
sample.input_ids, sample.loss_mask, sample.logprobs
sample.routed_experts, sample.sampling_maskCreates go round-robin over the servers, and each trajectory's URL names its
server, so no router or load balancer is involved. An SDK retry
(x-stainless-retry-count) gets the original call's reply rather than a second
sample.
Each trajectory is written once, when it ends (finish, idle TTL or graceful
shutdown): a document, plus sidecars for tokens (with the text they decode to
and each token's offset in it), routed experts and sampling masks. The format
is specified in docs/format.md, which is what any reader,
such as the viewer, implements.
cd skycap
uv sync --extra tokens
uv run pytestFormatting and lint are the repository's (bash format.sh from the root).