fix(sensevoice): preserve SentencePiece word boundaries - #3737
Merged
Merged
Conversation
Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SenseVoiceSmall.postimplementation.Head:
a57cc3f92a5064cdb4dd8f9417d2f03316ea77b3, based on55203a99956fb2b18f9c1ed9cc7cd78ce29a9a52. Three files only; the production change is two lines. Both commits are SSH-signed and DCO-signed-off; no history was rewritten.Reproduction
The first-token branch previously appended the token without stripping its boundary marker. It then assigned that marked string to
prev_word. Because the marker is non-ASCII, a following English subpiece failed the existing ASCII merge condition.Calling the real imported method, without constructing or loading a model:
Previously this returned words
["\u2581Hel", "lo", "world", "."]. It now returns["Hello", "world", "."], with the first word spanning[100, 300]ms. Subpiece intervals are combined using the existing rule: the first piece's start and last piece's end. The conversion from seconds to milliseconds and later word-boundary rules are unchanged.Related #3587 normalizes markers while aligning punctuation tokens; it does not normalize the raw word list or repair first-word subpiece joining. That existing defensive alignment code remains unchanged.
There is a second boundary case:
["Hello", "\u2581", "world"]previously became["Helloworld"], because skipping the standalone marker preserved the previous word. It now remains two words with their separate intervals. The current canonical SenseVoice postprocessor already handles both boundaries correctly; this aligns FunASR's bundled copy without an unrelated model refactor.Validation
torch==2.10.0+cpuandtorchaudio==2.10.0+cpu, using the current checkout as the import root:test_punc_model_none.py,test_timestamp_tools.py, andtest_sensevoice_tokenizer_special_tokens.py: 24 passed, 8 subtests passed.funasr,examples, andtests; two existing invalid-escape warnings informat5res.pyremain.git diff --checkand independent source review passed, with no actionable P1/P2 findings.An initial attempt to overlay NumPy versions onto a reused application environment did not pass: SciPy was incompatible with NumPy 1.26, and fresh subprocess imports lacked required dependencies. Those results are not counted as passing validation. The results above come from complete native CPU environments, without changing or skipping test assertions.
Boundaries
No checkpoint download, acoustic inference, GPU/MPS test, model-quality benchmark, or release was performed. Tests use synthetic token/time inputs and existing controlled unit fixtures. This fixes FunASR's bundled implementation; separately loaded external
remote_codecopies are not changed by this PR. No existing user issue is auto-closed.Hosted CI Snapshot
15c19598110268237a3e0ca9cdcd33c83a0b3f89. The merge tree exactly matches the validated tree0245827c4fa937605524367d60b908337f95f1ea; GitHub verifies the merge signature. Existing protection settings were not changed and no review was fabricated.The complete 38-case CPU selection passed again in each native NumPy environment immediately before merge; adjacent coverage again passed with 24 tests and 8 subtests. These checks do not claim acoustic, GPU, release, or production website validation.