Skip to content

feat: generalise hardcoded peephole optimisation families - #468

Draft
rkalis wants to merge 2 commits into
nextfrom
feat/dynamic-peephole-optimisations
Draft

rkalis wants to merge 2 commits into
nextfrom
feat/dynamic-peephole-optimisations

Conversation

@rkalis

@rkalis rkalis commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Follow-up to #458, which left a note in optimisations.ts that its hardcoded rule families "should become dynamic". This replaces the families that only differ in a stack depth or push with matcher functions, and generalises the in-place update to any expression.

Opened as a draft because of the two open decisions below.

Open decisions

1. For-loop update annotation in the BitAuth output. Opcodes that an optimisation keeps now keep their own source locations (see Optimiser engine below). For i = i + 1 loop updates the assignment's location was on the removed OP_NIP, so the annotation now reads >>> for-loop update (i + 1) instead of (i = i + 1). i++, i += 1 and i -= 3 still show the full statement, and the multi-line update in the NestedForWithDoWhile fixture now shows (i + 1) instead of an empty (). x = x + 1 is the most common form in the docs and tests (16 of 22 loops).

Restoring the old text means handing the location of a removed part to the opcode before it. But the removed drop can also be a scope or function cleanup: in compound_assign, that would give the final require check the whole function body's location and display it at the closing }. Options:

  • A (current): keep per-opcode locations and accept (i + 1).
  • B: only hand a removed part's location to the previous opcode when both are on the same line.
  • C: derive the for-loop update annotation from the loop header in the formatter, independent of optimisations.

2. <pick n> OP_ROT OP_SWAP OP_DROP → OP_SWAP never fires. The earlier OP_SWAP OP_DROP → OP_NIP rule always rewrites its tail first within the same pass, so this family (and the missing OP_15 rule mentioned in #458) has never had an effect. This PR generalises it as-is. Options: retarget it to the reachable <pick n> OP_ROT OP_NIP → OP_SWAP (same semantics, changes no current contract), or delete it.

Changes

Rules can be matcher functions (packages/utils/src/optimisations.ts)

The rule list holds [pattern, replacement] pairs and matcher functions, and went from 159 static rules to 94 static rules plus 3 matchers:

Matcher Replaces Now covers
updateInPlace: <pick n> <expression> <drop n+1> → <roll n> <expression> the 26 OP_1ADD/OP_1SUB rules from #458, OP_DUP OP_1ADD OP_NIP, OP_DUP OP_1SUB OP_NIP any depth (including above 16), any expression
moveDropBeforePush: <push> OP_NIP → OP_DROP <push> the 17 OP_n OP_NIP rules any push, including data and OP_1NEGATE
dropPickedCopy: <pick n> OP_ROT OP_SWAP OP_DROP → OP_SWAP 16 OP_n rules any depth (see open decision 2)

<pick n>, <roll n> and <drop n> are the forms left after the hardcoded stack op rules, e.g. OP_OVER, OP_ROT and OP_NIP for small depths.

The expression of an in-place update can be any sequence of opcodes that replace the items they pop with results that only depend on those items and the transaction, as long as it never consumes items below its own. A table of about 75 opcodes and their stack effects defines these (arithmetic, comparisons, hashes, splice, bitwise and shift ops, introspection, stack shuffles). The expression can also copy other variables. Copies of items below the updated variable are rewritten one shallower, because the roll removes it from the stack. This optimises e.g. total = total + tx.outputs[i].value inside a branch or loop.

It does not match when the expression:

  • copies the original variable again (e.g. x = x * x), since the expression might already have consumed its copy
  • contains any other opcode, e.g. OP_VERIFY, OP_IF, OP_ROLL, alt stack ops, OP_DEPTH or OP_INVOKE
  • is longer than 100 opcodes. This keeps the search linear: on a 30k-opcode worst case it would otherwise take seconds.

Optimiser engine (packages/utils/src/script.ts)

  • Rules match ASM tokens directly instead of regexes. For static rules the output, including all debug information, is identical to the regex engine: a differential fuzz of 200k random scripts with random metadata showed no differences apart from the range-end fix below.
  • A match is a list of consecutive parts, each with its own replacement. Static rules are a single part, which keeps their behaviour. The in-place update returns [pick → roll], one part per expression opcode, and [drop → nothing], so the opcodes it keeps keep their source locations and positions. Log data still adjusts against the whole match, because the roll changes the stack below the expression.
    Merging one location over a whole long match broke debugging: for x -= 3; require(x == 12), the final require's ip ended up one opcode early, so the SDK could no longer attribute the failed require. Every opcode also got the whole function body's location from the cleanup OP_NIP.
  • The inclusive end of a source tag or inline range inside a part that is removed entirely now moves to the opcode before it (instead of the one after), and tags whose opcodes were all removed are dropped. This removes a stale inline range in tuple_reassignment that attributed an unrelated OP_ROT to an inlined swap that was optimised away.

Fixtures

  • 18 generation fixtures and the BitAuth fixtures are updated.
  • The second commit regenerates the OverlappingScopeCleanup BitAuth fixture from fix: fix failing debug() when scope cleanup tags overlapped #463. It predates the OP_2 OP_PICK OP_2 OP_PICK → OP_3DUP OP_DROP rule from feat: add more peephole optimisations #458, and currently fails CI on next. The overlapping-tag test in bitauth-script.test.ts now uses the shifted indices of the same scope cleanups.
  • The release notes are unchanged: the existing line about new optimisations for loop counter updates and reassignments covers this.

Results

6 contracts get smaller, by 14 bytes in total. The cashc test contracts go from 3,157 to 3,143 bytes.

Contract Before After
complex_loop (outputSum = outputSum + tx.outputs[i].value in a loop branch) 64 61
for_loop_stack_items 35 32
compound_assign 14 12
global_function_in_control_flow 37 35
increment_decrement 12 10
tuple_modifiers 35 33

13 more contracts only change in debug information: the per-opcode locations of in-place updates.

Verification

  • yarn test (957 passed), yarn lint and yarn spellcheck pass.
  • New packages/utils/test/optimisations.test.ts covers each matcher, including cases that must not match. It checks VM equivalence on libauth's BCH 2026 VM with a stack of distinct numbers, so reading the wrong item changes the result. It also covers the debug information: kept locations, the final require ip, and tag ends.
  • VM fuzzing on libauth's BCH 2026 VM, through the full optimiseBytecode with 6 random stacks per script (the harness is not part of this PR):
    • About 116k random in-place updates (9 runs of about 13k unique scripts; random opcodes from the table, copies above, below and of the original variable, and opcodes that must end the expression): 0 mismatches, with about 34k successful runs through the rewritten copies.
    • About 158k random scripts built from all rule fragments (7 runs of about 22.6k unique scripts): 0 mismatches.
    • The only cases where the optimised script succeeds and the original fails are stack underflows (a copy of a variable that does not exist, which compiled code cannot contain) and the documented OP_NOT OP_IF and OP_NOT OP_NOT OP_VERIFY relaxations.
    • OP_ACTIVEBYTECODE, OP_UTXOBYTECODE and OP_OUTPOINTTXHASH were left out of the fuzz, because in libauth's test programs their values depend on the script bytes that the optimisation changes.
  • Performance: 23 ms instead of 20 ms for a realistic 5.8k-opcode contract. The worst case of 30k opcodes that could all be part of an expression takes 175 ms instead of 48 ms.

Possible follow-ups

  • OP_DUP OP_ROT OP_DROP, OP_OVER OP_ROT OP_DROP and OP_2 OP_PICK OP_ROT OP_DROP look like another depth family (<pick n> OP_ROT OP_DROP equals OP_NIP <pick n-1> for n ≥ 2, one opcode shorter). Not measured.
  • The fuzz harnesses could become a seeded test in the repo.

🤖 Generated with Claude Code

rkalis and others added 2 commits September 29, 2026 11:41
Replace the rule families that only differed in a stack depth or push with
matcher functions: an in-place update of a variable at any depth with any
pure expression (including copies of other variables), OP_NIP after any
push, and a dropped pick at any depth. Optimisation rules can now be either
[pattern, replacement] pairs or matchers that return their match as parts,
so opcodes a match keeps also keep their source locations and positions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The fixture from #463 predates the OP_2 OP_PICK OP_2 OP_PICK => OP_3DUP
OP_DROP optimisation from #458, which makes it fail on next. The
regenerated fixture also picks up the per-opcode source location of the
in-place loop counter update, and the overlapping tag test now uses the
shifted indices of the same scope cleanups.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Sep 29, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
cashscript Ready Ready Preview Sep 29, 2026 9:46am UTC

Request Review

@codecov

codecov Bot commented Sep 29, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.66667% with 11 lines in your changes missing coverage. Please review.
⚠️ Please upload report for BASE (next@4f157e2). Learn more about missing BASE report.

Files with missing lines Patch % Lines
packages/utils/src/script.ts 82.75% 6 Missing and 4 partials ⚠️
packages/utils/src/optimisations.ts 98.64% 1 Missing ⚠️
Additional details and impacted files
@@           Coverage Diff           @@
##             next     #468   +/-   ##
=======================================
  Coverage        ?   90.69%           
=======================================
  Files           ?       61           
  Lines           ?     5158           
  Branches        ?      979           
=======================================
  Hits            ?     4678           
  Misses          ?      367           
  Partials        ?      113           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

This branch was successfully deployed

1 active deployment
Preview — af30e11f Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant