Skip to content

fix(sensevoice): preserve SentencePiece word boundaries - #3737

Merged
LauraGPT merged 2 commits into
mainfrom
codex/fix-sensevoice-first-word-20260929
Sep 29, 2026
Merged

LauraGPT merged 2 commits into
mainfrom
codex/fix-sensevoice-first-word-20260929

Conversation

@LauraGPT

@LauraGPT LauraGPT commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Normalize the leading SentencePiece word-boundary marker in the first token handled by the bundled SenseVoiceSmall.post implementation.
  • Let the existing ASCII-subpiece joining logic handle the first word just as it handles later words, keeping words and their millisecond intervals paired.
  • Reset the previous-word state at a standalone boundary marker so separate English words are not joined across it.
  • Add 15 native-method regression cases and run them in the existing CPU NumPy compatibility workflow, including both pull-request and main-push path filters. The existing matrix versions are unchanged.

Head: a57cc3f92a5064cdb4dd8f9417d2f03316ea77b3, based on 55203a99956fb2b18f9c1ed9cc7cd78ce29a9a52. Three files only; the production change is two lines. Both commits are SSH-signed and DCO-signed-off; no history was rewritten.

Reproduction

The first-token branch previously appended the token without stripping its boundary marker. It then assigned that marked string to prev_word. Because the marker is non-ASCII, a following English subpiece failed the existing ASCII merge condition.

Calling the real imported method, without constructing or loading a model:

SenseVoiceSmall.post(None, [
    ["\u2581Hel", 0.1, 0.2],
    ["lo", 0.2, 0.3],
    ["\u2581world", 0.4, 0.6],
    [".", 0.6, 0.66],
])

Previously this returned words ["\u2581Hel", "lo", "world", "."]. It now returns ["Hello", "world", "."], with the first word spanning [100, 300] ms. Subpiece intervals are combined using the existing rule: the first piece's start and last piece's end. The conversion from seconds to milliseconds and later word-boundary rules are unchanged.

Related #3587 normalizes markers while aligning punctuation tokens; it does not normalize the raw word list or repair first-word subpiece joining. That existing defensive alignment code remains unchanged.

There is a second boundary case: ["Hello", "\u2581", "world"] previously became ["Helloworld"], because skipping the standalone marker preserved the previous word. It now remains two words with their separate intervals. The current canonical SenseVoice postprocessor already handles both boundaries correctly; this aligns FunASR's bundled copy without an unrelated model refactor.

Validation

  • Regression sequence: 5 failed / 7 passed before the first-word fix, then 12 passed. Adding the standalone-boundary cases produced 3 failed / 12 passed before the second fix; the final suite has 15 passed. The test imports and calls the actual class, not an AST copy or substituted implementation.
  • Python 3.12.3, matching torch==2.10.0+cpu and torchaudio==2.10.0+cpu, using the current checkout as the import root:
    • NumPy 1.26.4: 38 passed across the complete existing matrix test selection plus the new test file.
    • NumPy 2.4.0: 38 passed across the same selection.
    • Both native environments passed dependency checks. The 38 cases are repeated across environments, not 76 distinct tests.
  • Adjacent test_punc_model_none.py, test_timestamp_tools.py, and test_sensevoice_tokenizer_special_tokens.py: 24 passed, 8 subtests passed.
  • Direct integration of the real postprocessor output with VideoCaptioner's current word/sentence parser: both modes produced the expected text and interval bounds. This exercises parser methods, not audio recognition or the desktop UI.
  • The actual imported FunASR method matched the canonical SenseVoice postprocessor on 2,801 synthetic token streams over seven token values and lengths zero through four. Only the reference function was AST-extracted; the implementation under test was imported normally. This is a bounded postprocessing comparison, not model-output or acoustic equivalence.
  • New test file passes Black 24.4.0 with line length 100. Syntax compilation passed for 646 tracked/new Python files under funasr, examples, and tests; two existing invalid-escape warnings in format5res.py remain.
  • Workflow parsing confirms both event filters and the CPU test command include the regression. git diff --check and independent source review passed, with no actionable P1/P2 findings.

An initial attempt to overlay NumPy versions onto a reused application environment did not pass: SciPy was incompatible with NumPy 1.26, and fresh subprocess imports lacked required dependencies. Those results are not counted as passing validation. The results above come from complete native CPU environments, without changing or skipping test assertions.

Boundaries

No checkpoint download, acoustic inference, GPU/MPS test, model-quality benchmark, or release was performed. Tests use synthetic token/time inputs and existing controlled unit fixtures. This fixes FunASR's bundled implementation; separately loaded external remote_code copies are not changed by this PR. No existing user issue is auto-closed.

Hosted CI Snapshot

  • Final-head NumPy compatibility run: both NumPy matrix jobs succeeded.
  • Merged through the ordinary SHA-guarded GitHub merge endpoint as 15c19598110268237a3e0ca9cdcd33c83a0b3f89. The merge tree exactly matches the validated tree 0245827c4fa937605524367d60b908337f95f1ea; GitHub verifies the merge signature. Existing protection settings were not changed and no review was fabricated.
  • Exact merge-SHA main-push NumPy compatibility and API documentation workflows both succeeded.

The complete 38-case CPU selection passed again in each native NumPy environment immediately before merge; adjacent coverage again passed with 24 tests and 8 subtests. These checks do not claim acoustic, GPU, release, or production website validation.

Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
@LauraGPT LauraGPT changed the title fix(sensevoice): normalize first word before joining subpieces fix(sensevoice): preserve SentencePiece word boundaries Sep 29, 2026
@LauraGPT
LauraGPT merged commit 15c1959 into main Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant