Skip to content

persist_failed on very large repositories: staging DB dump stops at exactly 4 GiB (Windows, v0.11.0) #2399

Description

@GMorandi

Version

codebase-memory-mcp 0.11.0

Platform

Windows (x64)

Install channel

GitHub release archive / install.sh / install.ps1

Binary variant

standard

What happened, and what did you expect?

Summary

Indexing a very large repository (~51K files, >1 GB of source) in full mode completes all analysis successfully, but the final publish step fails with persist_failed. Monitoring the cache directory shows the staging temp file grows to exactly 4,294,967,296 bytes (2^32 = 4 GiB), stops growing, and the run is eventually rolled back. This strongly suggests a 32-bit size/offset limitation somewhere in the dump/publish path (Windows build).

Environment

  • codebase-memory-mcp 0.11.0 (latest release as of 2026-09-28), native Windows amd64 executable
  • OS: Windows Server cloud VM, 64 GB RAM (tool memory budget auto-set to 32 GB; observed worker peak only ~3 GB)
  • Filesystem: NTFS, C: drive — 33 GB free at failure time, so this is not a disk-space issue
  • Repository: ~51,190 files, mostly Java (many large generated DAO/SOAP classes of 250–730 KB)

Steps to reproduce

codebase-memory-mcp cli index_repository --repo-path <large-repo> --mode full --persistence true
  1. Extraction and semantic analysis run ~15 minutes and complete cleanly — daemon log shows index.supervisor.reap outcome=clean exit_code=0.
  2. The publish step then fails:
{"project":"GWH","status":"persist_failed","hint":"The validated staging database could not be published. Check free disk space and permissions on the cache directory; the previous index may have been rolled back."}

Evidence

Sampling ~/.cache/codebase-memory-mcp/ every 30 s during the run:

-rw-r--r-- 1 Administrator 197121 4294967296 ... GWH.db.stage.R1tcSN.tmp.96340.0000020e04424000

The temp file remains at exactly 4294967296 bytes across multiple samples spanning ~4 minutes, then disappears (rollback).

  • Reproducible: 3/3 consecutive attempts failed identically
  • Smaller repositories on the same machine index and publish fine (e.g. 395 files → 39 MB DB, status: indexed)
  • Stale .lock files from earlier killed attempts were removed before retrying — not the cause
  • For scale reference: after splitting the same codebase into per-directory projects, the resulting DBs sum to ~10.8 GB, so the single-project DB would have been well above 4 GiB

Expected behavior

Staging databases larger than 4 GiB publish successfully (e.g. chunked/streamed dump with 64-bit offsets).

Actual behavior

Once the staging database exceeds 4 GiB, the dump stalls at exactly 2^32 bytes, publish fails, and ~15 minutes of indexing work is rolled back.

Workaround

Index sub-directories as separate projects (keeping each DB under 4 GiB) and link them with --mode cross-repo-intelligence. This works, but loses cross-module SIMILAR_TO / SEMANTICALLY_RELATED edges and fragments the project list — not ideal for monorepos.

Secondary observation (happy to file separately if you prefer)

With 66 projects registered, list_projects exceeds 60 s via MCP (typical client timeout) and takes ~90 s via CLI. Per-project queries (search_graph, trace_path, get_architecture) stay sub-millisecond. Consider caching project stats or lazy enumeration for large project counts.

Reproduction

Reproduction

Code being indexed: The affected repo is private, but the failure is purely size-dependent, so any corpus whose full-mode index DB exceeds 4 GiB reproduces it. A deterministic synthetic corpus (mimics our real case: tens of thousands of near-identical generated Java classes):

# Generates 15,000 similar Java files (~1.5 GB source) -> full-mode DB > 4 GiB
mkdir -p repro/src/gen
{
  echo "package gen;"
  echo "class T {"
  for m in $(seq 1 1000); do
    echo "  public long method$m(long a, long b) { long c = a * $m + b; for (int i = 0; i < $m; i++) c += i * a; return c; }"
  done
  echo "}"
} > /tmp/template.java
for i in $(seq 0 14999); do cp /tmp/template.java repro/src/gen/Gen$i.java; done

(A large real-world public repo such as torvalds/linux — your own published benchmark at 75K files — likely also works, but the synthetic corpus above guarantees the >4 GiB threshold and needs no multi-GB clone.)

Exact command:

codebase-memory-mcp cli index_repository --repo-path C:/repro --mode full --persistence true

MCP equivalent (needs a client without a 60 s call timeout, or watch the cache dir after the call times out):

{"repo_path": "C:/repro", "mode": "full", "persistence": true}

What happened:

  1. Extraction/semantic analysis completes cleanly (~15 min on a 64 GB Windows VM; daemon log: index.supervisor.reap outcome=clean exit_code=0).
  2. In ~/.cache/codebase-memory-mcp/, the staging file …​.db.stage.XXXX.tmp.PID.…​ grows to exactly 4,294,967,296 bytes (2^32) and stops; sampled every 30 s it never grows past that value:
-rw-r--r-- 1 Administrator 197121 4294967296 Sep 28 12:10 GWH.db.stage.R1tcSN.tmp.96340.0000020e04424000
  1. After several minutes the command fails:
{"project":"GWH","status":"persist_failed","hint":"The validated staging database could not be published. Check free disk space and permissions on the cache directory; the previous index may have been rolled back."}

Environment: codebase-memory-mcp 0.11.0 native Windows amd64; NTFS with 33 GB free (not a space issue); stale locks ruled out; reproducible 3/3 runs. The same machine indexes smaller repos fine.

What should have happened:

{"project":"GWH","status":"indexed","nodes":...,"edges":...}

i.e. staging DBs larger than 4 GiB publish successfully (chunked/streamed dump or 64-bit offsets), instead of stalling at exactly 2^32 bytes and rolling back ~15 minutes of work.

Logs


Diagnostics trajectory (memory / performance / leak issues)


Project scale (if relevant)

No response

Confirmations

  • I searched existing issues and this is not a duplicate.
  • My reproduction uses shareable code (a dummy snippet or a public OSS repository), not proprietary code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingparsing/qualityGraph extraction bugs, false positives, missing edgesstability/performanceServer crashes, OOM, hangs, high CPU/memorywindowsWindows-specific issues

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions