RSR-27B: Recursive Self-Rewrite (RSR) Model

RSR-27B is Qwen-3.8-27B finetuned with Recursive Self-Rewrite (RSR) trajectories. We did additional RL training to stabalize the model on hard terminal tasks (Details coming soon).

Sample trajectories

3 SWR tasks that base Qwen3.8-27B passed under continue-until-timeout. For each: the expert trajectory, the runbook the base model wrote from it, and a passing rewrite (base Qwen3.8-27B as plain Terminus-2, runbook kept private). Browse them side by side in the interactive trajectory viewer.

Expert vs. passing rewrite per task: commands by kind, steps, and agent time

Task Domain Steps, expert โ†’ rewrite Agent time, expert โ†’ rewrite Runbook replays passed
tree_sitter_markdown_inline__05 Data, query & formal (tree-sitter) 64 โ†’ 34 2.0 h โ†’ 22 min 4 of 4
cmake_generated_header_config__06 Systems & infrastructure (CMake) 59 โ†’ 27 2.0 h โ†’ 21 min 4 of 4
openfoam_scalar_transport_boundary_probe__02 Physical science & simulation (OpenFOAM) 56 โ†’ 23 1.9 h โ†’ 28 min 4 of 4
tree_sitter_markdown_inline__05: overview and step lists

Open this task side by side

Data, query & formal (tree-sitter). The expert never marked the task complete: it kept working until the 2-hour limit ended the run, and the work it left in place passed. Runbook: runbook_01, checked by the leak judge; 4 of 4 replays passed, in 13, 15, 15, 34 steps, and the one shown took 34.

Shortened runbook (a brief summary, not the actual runbook used): A short plan for rebuilding the tree-sitter Markdown demo: learn what the reference outputs look like, write a script that reproduces them, and check that it generalizes beyond the public examples.

Expert โ€” Base Qwen3.8-27B ยท continue-until-timeout: 64 steps, 153 commands

  • Step 1: First, let me explore the directory structure to understand what's available - observations, tools, and any existing files. โ€” ls -la /app/
  • Step 2: Read the manifest.json, extract the reference outputs, read the feedback tool, and the README to understand what's expected. โ€” cat /app/observations/manifest.json
  • Step 3: Read all the extracted files to understand the expected output format, parameters, and behavior. โ€” cat /tmp/obs/public_manifest.json
  • Step 4: Read the native files (generate.log, parse.txt, parse_metrics.json) and the oracle pretty output for both public scenarios to understand the full pict โ€” cat /tmp/obs/public_1/native/generate.log
  • Step 5: Read public_2's native files to compare with public_1 and understand how parameters affect the output. โ€” cat /tmp/obs/public_2/native/generate.log
  • Step 6: Let me write a quick Python script to verify the signature computation, then create the initial reconstruct.py and test it with the feedback tool. โ€” python3 -c "
  • Step 7: Let me check if tree-sitter is available, look for any input files, and understand the mechanism better. Then I'll create the initial reconstruct.py. โ€” which tree-sitter 2>/dev/null; tree-sitter --version 2>/dev/null; echo "exit: $?"
  • Step 8: Let me explore the tree-sitter CLI capabilities, check for markdown grammar, and understand the parse tree structure better. Then I'll create the init โ€” tree-sitter --help 2>&1 | head -30
  • โ€ฆ 53 more steps โ€ฆ
  • Step 62: Let me look at the tree-sitter-markdown-inline grammar to understand how it works, and try to find the input bundle or any other clues about the param โ€” cat /root/github/ts-md3/package/tree-sitter-markdown-inline/grammar.js | head -100
  • Step 63: Let me create a comprehensive test to understand the parameter-to-input mapping, and then build a robust reconstruct.py that can handle hidden scenari โ€” python3 << 'PYEOF'
  • Step 64: Let me create a comprehensive test to understand the parameter-to-input mapping, and then build a robust reconstruct.py that can handle hidden scenari โ€” python3 << 'PYEOF'

Rewrite โ€” Base Qwen3.8-27B ยท plain Terminus-2, runbook kept private: 34 steps, 78 commands

  • Step 1: Step 1: Explore the directory structure to understand what's available - observations, tools, and any existing files. Step 2: Extract and examine the โ€” ls -la /app/
  • Step 2: 1. Read the manifest.json and README.md 2. Extract the reference_outputs.tar.gz to examine oracle outputs 3. Read the feedback tool to understand how โ€” cat /app/observations/manifest.json
  • Step 3: Read the oracle outputs, native files, and manifest to understand the exact structure and expected behavior. โ€” cat /tmp/obs/public_1/oracle_output.json
  • Step 4: 1. Read the native output files (generate.log, parse.txt, parse_metrics.json) 2. Read the feedback tool 3. Read the scenario.json and signature.json f โ€” cat /tmp/obs/public_1/native/generate.log
  • Step 5: 1. Read the feedback tool to understand signature computation 2. Verify the signature computation by computing sha256 of the summary object 3. Check i โ€” cat /app/tools/feedback
  • Step 6: 1. Analyze the native files to understand how numbers and keywords are extracted 2. Check the artifact_manifest.json for additional context 3. Check t โ€” python3 -c "
  • Step 7: 1. Analyze keyword extraction from native files 2. Check if tree-sitter markdown grammar is available 3. Check what markdown input produces the observ โ€” python3 -c "
  • Step 8: 1. Check for tree-sitter markdown grammar availability 2. Test tree-sitter parse with a sample markdown 3. Analyze the keyword extraction rules more c โ€” find / -name '*.node' 2>/dev/null | head -20
  • โ€ฆ 23 more steps โ€ฆ
  • Step 32: Run final checks: verify files exist, run feedback one more time, and test the reconstruct.py with both public scenarios. โ€” ls -la /app/workspace/
  • Step 33: Task is complete. All deliverables are in place and validated. (marks done)
  • Step 34: Confirm task completion. (marks done)
cmake_generated_header_config__06: overview and step lists

Open this task side by side

Systems & infrastructure (CMake). The expert never marked the task complete: it kept working until the 2-hour limit ended the run, and the work it left in place passed. Runbook: runbook_02, checked by the leak judge; 4 of 4 replays passed, in 24, 27, 27, 29 steps, and the one shown took 27.

Shortened runbook (a brief summary, not the actual runbook used): A short plan for reproducing the CMake configuration demo: work out which outputs change with the scenario parameters, write a script that reproduces them, and check that it generalizes beyond the public examples.

Expert โ€” Base Qwen3.8-27B ยท continue-until-timeout: 59 steps, 94 commands

  • Step 1: First, explore the directory structure to understand what's available - observations, tools, and any existing files. โ€” ls -la /app/
  • Step 2: Read the manifest.json, README.md, and extract the reference outputs to understand the expected behavior. โ€” cat /app/observations/manifest.json
  • Step 3: Extract the reference_outputs.tar.gz to examine the public manifest and oracle outputs, understanding the expected output format. โ€” mkdir -p /tmp/ref && cd /tmp/ref && tar xzf /app/observations/reference_outputs.tar.gz &&
  • Step 4: Examine the extracted public_1 and public_2 directories, and get the full JSON structure of the public manifest to understand all fields. โ€” find /tmp/ref/public_1 -type f | head -30
  • Step 5: Let me examine the oracle_output.json files directly and compare the two public scenarios to understand what changes with parameters. โ€” cat /tmp/ref/public_1/oracle_output.json | python3 -m json.tool | head -100
  • Step 6: Compare the two oracle outputs, look at the native CMakeCache.txt files to see how parameters are embedded, and understand the full output structure. โ€” python3 -c "
  • Step 7: Compare the two oracle outputs in detail to understand what changes with parameters, and examine the native files more closely. โ€” python3 << 'EOF'
  • Step 8: Examine the native CMakeCache.txt and CMakeConfigureLog.yaml files in detail to understand the exact format and how parameters affect them. โ€” cat /tmp/ref/public_1/native/build/CMakeCache.txt
  • โ€ฆ 48 more steps โ€ฆ
  • Step 57: First, let me look at the current reconstruct.py, the workspace files, observations, and the feedback tool to understand the full picture. โ€” cat /app/workspace/reconstruct.py
  • Step 58: (reply could not be parsed; the harness asked for valid JSON)
  • Step 59: First, let me look at the current reconstruct.py, the workspace files, observations, and the feedback tool to understand the full picture. โ€” cat /app/workspace/reconstruct.py

Rewrite โ€” Base Qwen3.8-27B ยท plain Terminus-2, runbook kept private: 27 steps, 58 commands

  • Step 1: 1. List contents of /app/observations and /app/tools 2. Extract the reference outputs tarball 3. Examine the public manifest and oracle outputs to und โ€” ls -la /app/observations/
  • Step 2: 1. Look at the public manifest structure 2. Examine the oracle_output.json for both public scenarios 3. Compare the two to understand what changes bet โ€” cat /tmp/ref/public_manifest.json | python3 -m json.tool | head -80
  • Step 3: 1. Look at the scenario.json files for both public scenarios 2. Compare the oracle_output.json structures 3. Look at the native files that differ betw โ€” cat /tmp/ref/public_1/scenario.json
  • Step 4: 1. Compare the two oracle outputs to find which files differ 2. Examine the CMakeCache.txt differences (likely contains parameter values) 3. Look at t โ€” python3 -c "
  • Step 5: 1. Compare the CMakeCache.txt entries between the two oracle outputs 2. Compare the CMakeConfigureLog.yaml entries 3. Look at the native CMakeCache.tx โ€” python3 -c "
  • Step 6: 1. Look at the native CMakeCache.txt to see the parameter entries 2. Look at CMakeConfigureLog.yaml to understand the cmtc_ IDs 3. Understand the full โ€” grep -n 'SILR_PARAM\|CMAKE_BUILD_TYPE' /tmp/ref/public_1/native/build/CMakeCache.txt
  • Step 7: 1. Search for cmtc_ in the yaml file more broadly 2. Look at the full entry structure for CMakeCache.txt and CMakeConfigureLog.yaml 3. Understand how โ€” grep -i 'cmtc' /tmp/ref/public_1/native/build/CMakeFiles/CMakeConfigureLog.yaml | head -10
  • Step 8: 1. See the full CMakeCache.txt entry structure (all fields) 2. See the full CMakeConfigureLog.yaml entry structure 3. Understand the CMAKE_BUILD_TYPE โ€” python3 -c "
  • โ€ฆ 16 more steps โ€ฆ
  • Step 25: 1. Run one more comprehensive test with a hidden scenario 2. Verify the full output structure matches expectations 3. Confirm the feedback tool passes โ€” echo '{"parameters":{"build_type":5.5,"configured_version":4.4,"consumer_project":3.3},"sc
  • Step 26: The task is complete. The solution correctly handles both public and hidden scenarios. (marks done)
  • Step 27: Task is complete. Marking as done. (marks done)
openfoam_scalar_transport_boundary_probe__02: overview and step lists

Open this task side by side

Physical science & simulation (OpenFOAM). The expert never marked the task complete: it kept working until the 2-hour limit ended the run, and the work it left in place passed. Runbook: runbook_01, checked by the leak judge; 4 of 4 replays passed, in 23, 23, 33, 51 steps, and the one shown took 23.

Shortened runbook (a brief summary, not the actual runbook used): A short plan for reproducing the OpenFOAM boundary-probe demo: work out which results depend on the scenario parameters, write a script that reproduces them, and check that it generalizes beyond the public examples.

Expert โ€” Base Qwen3.8-27B ยท continue-until-timeout: 56 steps, 72 commands

  • Step 1: First, explore the directory structure to understand what's available - observations, tools, and any existing files. โ€” ls -la /app/
  • Step 2: Read the manifest.json, README.md, feedback tool, and extract the reference outputs to understand the expected behavior. โ€” cat /app/observations/manifest.json
  • Step 3: Extract the reference_outputs.tar.gz to examine the expected output format and understand the oracle behavior. โ€” mkdir -p /tmp/ref && cd /tmp/ref && tar xzf /app/observations/reference_outputs.tar.gz &&
  • Step 4: View the full public_manifest.json content, and also check what other files are in the extracted tar. โ€” ls -la /tmp/ref/
  • Step 5: Save the full JSON to a file and examine its structure programmatically. โ€” python3 -c "
  • Step 6: Examine the structure of parameter_effect and summary for both scenarios. โ€” python3 -c "
  • Step 7: Examine the summary structure for both scenarios in detail. โ€” python3 -c "
  • Step 8: Save the full oracle output to a file and examine the top-level structure of the summary. โ€” python3 -c "
  • โ€ฆ 45 more steps โ€ฆ
  • Step 54: Create a basic reconstruct.py that uses the oracle data directly to see if the feedback tool accepts it. Then I'll work on understanding the formulas. โ€” python3 << 'PYEOF'
  • Step 55: Create a basic reconstruct.py that uses the oracle data directly to see if the feedback tool accepts it. Then I'll work on understanding the formulas. โ€” python3 << 'PYEOF'
  • Step 56: Create a basic reconstruct.py that uses the oracle data directly to see if the feedback tool accepts it. Then I'll work on understanding the formulas. โ€” python3 << 'PYEOF'

Rewrite โ€” Base Qwen3.8-27B ยท plain Terminus-2, runbook kept private: 23 steps, 51 commands

  • Step 1: 1. List files in /app/observations and /app/tools 2. Extract the reference_outputs.tar.gz to understand the output structure 3. Examine the public man โ€” ls -la /app/observations/
  • Step 2: 1. Read the manifest.json to understand the scenario structure 2. Extract reference_outputs.tar.gz to see the oracle outputs 3. Read the README.md in โ€” cat /app/observations/manifest.json
  • Step 3: 1. Look at the public_manifest.json 2. Examine the oracle_output.json structure for public_1 3. Compare the two scenarios to understand what changes โ€” python3 -c "import json; d=json.load(open('/tmp/ref/public_manifest.json')); print(json.du
  • Step 4: 1. Get the full structure of oracle_output.json including all keys 2. Compare the two public scenarios to identify what changes 3. Look at the signatu โ€” python3 -c "
  • Step 5: 1. Read the feedback tool to understand how matching works 2. Compare all 24 output files between the two scenarios to find what changes 3. Look at sc โ€” cat /app/tools/feedback
  • Step 6: 1. Compare all 24 output files between the two scenarios to identify which ones change 2. For files that change, identify the pattern of change 3. Bui โ€” python3 -c "
  • Step 7: 1. Compare the blockMesh.log entries between the two scenarios in detail 2. Look at the actual native blockMesh.log files to understand the content 3. โ€” python3 -c "
  • Step 8: 1. Look at the blockMeshDict to understand the mesh parameters 2. Compare the two blockMesh.log native files 3. Analyze the formula for the changing v โ€” cat /tmp/ref/public_1/native/case/system/blockMeshDict
  • โ€ฆ 12 more steps โ€ฆ
  • Step 21: 1. Re-examine the blockMesh.log more carefully for any other parameter-dependent values 2. Check if the mesh statistics in the log should change with โ€” python3 -c "
  • Step 22: The solution is complete. Let me do one final verification and mark the task as complete. โ€” python3 /app/tools/feedback (marks done)
  • Step 23: Task is complete. Both public scenarios pass and the solution generalizes to hidden scenarios. (marks done)
Downloads last month
21
Safetensors
Model size
28B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for IntelligenceLab/RSR-27B

Quantizations
1 model

Space using IntelligenceLab/RSR-27B 1