Skip to content

Commit ecc4814

Browse files
authored
Merge pull request #21 from weiwch/main
Add FrontierCS algorithmic-problem-solving skill
2 parents 044b56b + 177cdf6 commit ecc4814

15 files changed

Lines changed: 8880 additions & 0 deletions

File tree

Lines changed: 80 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,80 @@
1+
# Algorithmic Problem-Solving Recovery
2+
3+
`algorithmic-problem-solving` is an evidence-driven recovery system for algorithmic, competitive-programming, interactive, online, and scored optimization tasks. It is not a template collection or a replacement for the first focused solution attempt. Once that attempt fails to achieve a verified full result, the skill preserves trusted work, identifies the failing layer, and selects the narrowest justified recovery route.
4+
5+
```text
6+
reproduce the failure -> recover the operational contract
7+
-> classify the failing layer -> run the selected recovery route
8+
-> validate a challenger independently -> promote the best legal champion
9+
-> escalate structurally when local improvement has stalled
10+
```
11+
12+
## Activation and routing
13+
14+
The root `SKILL.md` is the recovery entry point. It activates after few non-full submissions, after the first non-full candidate for an interactive or scored heuristic/optimization task, before a non-full final delivery, or when evidence shows a correctness, resource, evaluator, protocol, or policy failure. It does not apply before a fresh problem's first focused attempt, and it stops after a verified full pass or maximum score.
15+
16+
Sub-skills are targeted internal routes. The root selects them from the observed failure rather than loading every branch at once. Mandatory or paired routes remain additive when the task requires them.
17+
18+
## Codex adaptation
19+
20+
The repository does not ship product-specific `agents/openai.yaml` files, but a Codex adapter should allow implicit invocation only for the root router. Add or merge this policy into `algorithmic-problem-solving/agents/openai.yaml`:
21+
22+
```yaml
23+
policy:
24+
allow_implicit_invocation: true
25+
```
26+
27+
Every sub-skill should disable implicit invocation so Codex cannot bypass the root diagnosis and routing logic. For example, `algorithmic-problem-solving/sub-skills/interactive-problem-solving/agents/openai.yaml` should contain:
28+
29+
```yaml
30+
policy:
31+
allow_implicit_invocation: false
32+
```
33+
34+
Apply the same `false` policy to every other sub-skill. The snippets show only the invocation policy; no `default_prompt` is required for this routing design. These optional adapter files are intentionally omitted from the runtime structure below.
35+
36+
## Structure
37+
38+
```text
39+
algorithmic-problem-solving/
40+
├── SKILL.md # Recovery router, escalation gates, artifact discipline, and final release rules
41+
├── references/
42+
│ ├── heuristic-search.md # Quantified search budgets, representations, incremental evaluation, neighborhoods, and optimizers
43+
│ └── technique-selection.md # Algorithm-family selection after the model and resource envelope are trusted
44+
└── sub-skills/
45+
├── checker-and-local-evaluation/
46+
│ └── SKILL.md # Independent checkers, scorers, interactors, simulators, generators, and local evaluation
47+
├── contest-solver-engineering/
48+
│ └── SKILL.md # Toolchain, numeric, memory, runtime, I/O, randomness, and implementation recovery
49+
├── interactive-problem-solving/
50+
│ └── SKILL.md # Protocol modeling, hidden hypotheses, information-gaining queries, and transcript validation
51+
├── model-and-route-algorithms/
52+
│ └── SKILL.md # Contract/model repair, proof obligations, feasibility analysis, and solution-class selection
53+
├── plateau-escape/
54+
│ └── SKILL.md # Independent structural review, method research, executable challengers, and evidence-gated promotion
55+
├── reactive-online-decision-problem-solving/
56+
│ └── SKILL.md # Estimation, planning, exploration, feedback updates, and risk-aware sequential decisions
57+
├── testlib-cpp-judging/
58+
│ ├── SKILL.md # C++ checkers, validators, deterministic generators, interactors, and local judging flow
59+
│ ├── references/
60+
│ │ ├── testlib-usage.md # Testlib roles, APIs, verdicts, templates, and command-line contracts
61+
│ │ └── troubleshooting.md # Compilation, arguments, strict input, status, reproducibility, and protocol diagnostics
62+
│ └── scripts/
63+
│ └── testlib.h # Bundled single-header Testlib dependency
64+
└── validation-and-experiments/
65+
└── SKILL.md # Falsifying tests, independent oracles, paired comparisons, holdouts, and release gates
66+
```
67+
68+
## Operating principles
69+
70+
- **Recover the real contract first.** Read the statement and relevant executable artifacts, separate legality from objective and displayed score, and test any dependency on conflicting interpretations.
71+
- **Classify before editing.** Distinguish model/proof errors, route-selection errors, implementation failures, evaluator uncertainty, weak experimental evidence, search-mechanics problems, information-acquisition failures, and reward-bearing sequential decisions.
72+
- **Select methods from proven premises.** Use the technique catalog only after the model and feasibility envelope are trusted. Every reduction, optimized recurrence, advanced structure, or incomplete route needs an explicit proof or falsification target.
73+
- **Quantify scored search.** Design the representation, invariants, incremental evaluator, reachable neighborhoods, and useful-event rate before choosing an optimizer. Keep `current_state` separate from `best_valid_state`.
74+
- **Build local evaluation only when it is diagnostic.** A missing evaluator or a first non-full candidate is not sufficient by itself. Construct one when concrete legality, score, protocol, replay, or official/local disagreement makes it useful. Structural plateau recovery has its own stricter evaluator gate.
75+
- **Separate interactive and reactive work.** Interactive routing asks what information to acquire; reactive routing chooses reward-bearing actions whose live feedback changes later decisions. Offline repeated evaluation is neither by itself.
76+
- **Preserve artifact roles.** A `fallback` is the simplest guaranteed-valid emergency output, a `challenger` is experimental, a `champion` is the best independently validated legal artifact, and a `baseline` is an external evaluation reference.
77+
- **Promote mechanically from evidence.** A challenger replaces the champion only when it remains legal, respects resource and protocol limits, and improves comparable correctness or scoring evidence.
78+
- **Escalate structure instead of tuning indefinitely.** Once the root's plateau or severe-gap gate is met, freeze the champion, run isolated trajectory review and method research, implement the leading structural routes, and bind adoption or rejection to reproducible results.
79+
80+
Validation starts from the smallest faithful reproducer and expands only as needed through boundary cases, tiny brute-force oracles, differential and metamorphic tests, evaluator self-tests, paired seeds, holdouts, resource probes, and clean release runs. Only experiments that were actually executed count as evidence.

0 commit comments

Comments
 (0)