Repository navigation
After a vendored-to-hosted takeover is interrupted once its commit journal is written, the next rollback replays the journal, then exits 0 having restored nothing, and the project stays hosted-patched (remove says "No patch found") #1241
Description
Activity
- addedbugSomething isn't workingSomething isn't workingbughuntFound by a scheduled package-manager bug-hunt agentFound by a scheduled package-manager bug-hunt agentpm:yarn-berryYarn Berry (2+)Yarn Berry (2+)
on Oct 9, 2026 - added a commit that references this issue
on Oct 9, 2026 mikolalysenko commented
on Oct 9, 2026 CollaboratorAuthorMore actions[agent] Triage: priority:p1 (yarn berry is npm family; the defect is in shared CLI code:
rollback.rsandremove.rsbuildhosted_inventorybefore the apply lock replays the commit journal). Not a duplicate: related to #1157 (same interrupted-takeover window, but there the journal never records the deferred artifact removals) and #809. No open PR covers it.
Generated by Claude Code
mikolalysenko commented
on Oct 9, 2026 CollaboratorAuthorMore actions[agent] Re-checked by the scheduled Yarn Berry (2+) bug-hunt routine (ledger #305). This was closed as completed, but no commit references it, and it still reproduces on main
a80b89e(yarn 4.18.1, node-modules, Linux).rollback.rsstill runshosted_inventory(:913) beforeacquire_or_emit(:1011). The reverse takeover direction (hosted→vendored) is broken by the same ordering too; see the second table.The repro is now deterministic, with no timing loop. An
LD_PRELOADshim SIGKILLs the process whenrename()targets a given path, so the kill lands after.socket/vendor/.commit-journal.jsonis durable:// shim.c: gcc -shared -fPIC -o shim.so shim.c -ldl #define _GNU_SOURCE #include <dlfcn.h> #include <string.h> #include <stdlib.h> #include <signal.h> #include <unistd.h> static int hit(const char *n){ const char *t=getenv("KILL_ON"); if(!t||!n) return 0; size_t a=strlen(n),b=strlen(t); return a>=b && !strcmp(n+a-b,t); } int rename(const char *o,const char *n){ if(hit(n)) kill(getpid(),SIGKILL); int(*r)(const char*,const char*)=dlsym(RTLD_NEXT,"rename"); return r(o,n); } int renameat2(int a,const char*o,int b,const char*n,unsigned f){ if(hit(n)) kill(getpid(),SIGKILL); int(*r)(int,const char*,int,const char*,unsigned)=dlsym(RTLD_NEXT,"renameat2"); return r(a,o,b,n,f); }
# committed project (left-pad@1.3.0 + uuid@9.0.1); mock API as in the issue body LD_PRELOAD=./shim.so KILL_ON=/package.json socket-patch scan --mode hosted $API # vendored→hosted, killed (137) LD_PRELOAD=./shim.so KILL_ON=/yarn.lock socket-patch scan --mode vendored $API # hosted→vendored, killed (137) socket-patch rollback $API; echo $?
Vendored→hosted takeover (this issue):
kill point first rollbackproject afterwards second rollbackrename→package.json(no file replaced yet)exit 0, rolledBack: 0,hosted.reverted: []hosted-pinned (the replay finished the commit) exit 0, restored rename→yarn.lock(package.jsonalready replaced)exit 1 partial_failure,hosted.failed: [hosted_wiring_contested](the pre-lock discovery read the half-committed pair)hosted-pinned (the same lock's replay finished the commit) exit 0, restored Hosted→vendored takeover (
scan --mode vendoredover a hosted-pinned project). This is the same pre-lock ordering, seen from the other side:kill point first rollbackjournal / project afterwards rename→package.jsonreplays, prints Reverted vendoring for …×2, thenError: Cannot restore pkg:npm/left-pad@1.3.0 … no hosted wiring … was found in yarn.lock(and the same for uuid), exit 1. The pre-lock list held hosted pins the replay had already removedreplayed; project unpatched. The error and exit code are wrong rename→yarn.lockrefuses before taking the lock: yarn.lock wire(s) Socket-hosted patches that cannot be attributed … (patched_ref_invalid: … orphaned: no package.json resolutions entry …), exit 1.remove <purl>gives the same errorjournal never replayed: package.jsonholds the vendoredfile:resolutions whileyarn.lockstill holds the hosted entriesrename→state.json(both files replaced)refuses before taking the lock: Lockfiles still reference .socket/vendor/ artifacts but the vendor ledger is missing — restore .socket/vendor/state.json from version control, exit 1. That file was never committed (the project was hosted), and the ledger is in the pending journal.removegivesManifest not foundjournal never replayed; listreports "No patches in this project" (exit 0, #809)If the user follows the printed
git checkout -- yarn.lock package.jsonremedy in the last two rows, the nextrollbacktakes the lock and replays the journal over their checkout. It then reverts the vendoring and still fails with the staleCannot restore …errors (exit 1). The end state is unpatched.I ran every row at least twice with the same result.
remove 11111111-…after the vendored→hosted kills exits 1 withNo patch found matching identifier. The fix suggested in the issue body (run discovery and the pre-lock refusals afteracquire_or_emitreplays the journal, or redo them once the lock is held) covers both directions. After a full unwind, left-pad's entry is byte-exact; the only leftover diff is uuid'sbin: ./dist/bin/uuid, which is #1131.
Generated by Claude Code
[agent] Found by the scheduled Yarn Berry (2+) bug-hunt routine (ledger #305).
Summary
A vendored→hosted takeover (
scan --mode hostedover a vendored project, #1039) commits through the roll-forward journal.socket/vendor/.commit-journal.json. Suppose the process dies after the journal is durable but before any file is replaced, androllbackis the next command. Then:rollbackdiscovers the hosted pins before it takes the apply lock (crates/socket-patch-cli/src/commands/rollback.rs:908-915,hosted_inventory). At that momentyarn.lockandpackage.jsonstill hold the vendoredfile:wiring, so it finds no hosted pins.acquire_or_emit(rollback.rs:1008) takes the lock, andapply_lock::acquire→recover_group_commitreplays the journal. The hostedresolutionspins and re-keyed lock entries land, and.socket/vendor/state.jsonis deleted.rollbackexits 0 and prints nothing at all (human mode). JSON reportsstatus: success,rolledBack: 0andhosted.reverted: []. The project ends up hosted-pinned: every freshyarn install --immutableinstalls the patched tarball. A secondrollbackdoes restore it.remove <uuid>andremove <purl>from the same state replay the journal too, then fail withError: No patch found matching identifier: …(exit 1), although the patch is pinned in the lock they just rewrote.remove.rshas the same order:hosted_inventoryat:367, lock at:406.Impact
The user asked to remove every patch, got exit 0 and no error, and the project still installs Socket's patched bytes. CI or scripted unwinds (
rollback && git commit) commit a hosted-pinned lock. The window is narrow (a crash, SIGKILL or OOM between the journal write and the first file replace), but this is exactly the case the journal exists for. The CLI contract promises that the next locked command "replays the journal before reading anything … so a locked command never observes a half-committed run" (CLI_CONTRACT.md, "Vendored group commit"), and that "a commit interrupted after its journal was written is finished by the next command that takes the apply lock" ("Staged takeover").The cause is in shared CLI code, not in the yarn rewriter, so every ecosystem with a vendored→hosted takeover should be affected. I proved it only with real Yarn Berry.
Repro (real yarn 4.18.1 / 4.0.2, release build, Linux)
This uses a local mock of the patch API (
/v0/orgs/org/patches/{batch,package,view,by-package}plus/artifacts/<uuid>/<name>-<ver>.tgzwith ayarn-berry-zipyarnBerry10c0artifact), as in #1157. No failpoint build is needed: a realSIGKILLtimed into the commit window hits it within a few tries.Output (the same on every hit):
Apart from the token warning,
rollbackprinted nothing. A fresh checkout'syarn install --immutable(cold cache) then installs/* SOCKET-PATCHED … */left-pad. Runningrollbacka second time printsRestored pkg:npm/is-number@7.0.0 to its upstream registry entry(and the same for left-pad), andyarn.lockandpackage.jsonare byte-identical to the pre-vendor commit.State at the kill:
git statusshows only?? .socket/apply.lockand?? .socket/vendor/.commit-journal.json. The journal namespackage.json,yarn.lock(both withafterbytes) and.socket/vendor/state.json(deletion).From a copy of the same post-kill snapshot:
rollbackrollback --jsonsuccess,rolledBack: 0,hosted.reverted: []remove <uuid>/remove pkg:npm/left-pad@1.3.0No patch found matching identifiervendor --revertvendor --revertdoesn't unwind hosted).socket-stage-.commit-journal.json-*exists)Reverted vendoring for …×2Expected vs actual
rollback(andremove) plan against the project as it is after the journal replay. CLI_CONTRACT says a locked command replays the journal "before reading anything", so the replayed hosted pins are restored to the registry entries and the run reports them. Concretely: run the hosted-pin discovery (hosted_inventory) and the truly-empty probes afteracquire_or_emit, or re-run them once the lock is held when a journal was replayed.removerefuses a patch that is pinned.Matrix (Linux, Node 22, main
f3c6313)rollbackafter the hitmacOS and Windows weren't probed; this is OS-independent CLI ordering. Release 4.0.0 isn't affected: it has no journaled takeover. The first bad commit is probably 823810a (#1039, which made the takeover a journaled group commit); not bisected.
Suspect code
crates/socket-patch-cli/src/commands/rollback.rs:908-917:hosted_inventoryand the truly-empty decision run before the lock. The comment says only cheap existence probes run pre-lock, but the hosted pins are the leg's work list.crates/socket-patch-cli/src/commands/rollback.rs:1008:acquire_or_emit→socket_patch_core::patch::apply_lock::acquire→recover_group_commit(crates/socket-patch-core/src/patch/apply_lock.rs:576).crates/socket-patch-cli/src/commands/remove.rs:367/:406: the same order.Related but different: #1157 (the same crash window leaves the vendored artifact directory behind; there a
scanwas the next command) and #809 (lock-free readers see the pending journal).