Skip to content

fix: assign sentence speakers by total overlap without double counting - #3733

Merged
LauraGPT merged 1 commit into
mainfrom
codex/speaker-overlap-accounting-20260928
Sep 28, 2026
Merged

LauraGPT merged 1 commit into
mainfrom
codex/speaker-overlap-accounting-20260928

Conversation

@LauraGPT

Copy link
Copy Markdown
Collaborator

Problem

The shared CAM++ speaker assignment helper counts a newly leading segment's overlap twice and only accumulates segments for the currently leading speaker. For a 0-10 second sentence with speaker 0 at 0-4s and speaker 1 at 4-10s, the current real module returns speaker 0 despite speaker 1 covering 6 seconds.

Change

  • Accumulate each positive, sentence-clipped segment overlap once per speaker; select the greatest total after all segments.
  • Resolve equal totals deterministically using the first positively overlapping speaker. This explicitly defines ties rather than promising equivalence with every result of the old faulty algorithm.
  • Preserve millisecond/second conversion, integer labels, no-overlap fallback 0, in-place sentence/list identity, and all other fields.
  • Add 32 regression cases and wire them into both events and the test/format steps of the existing speaker-adapter workflow.

The normal caller supplies postprocessed diarization intervals. This does not add union/deduplication semantics for arbitrary overlapping or duplicate intervals.

Validation

  • Real-module RED: the new 4s/6s test failed on exact main da12182; GREEN after the fix.
  • 32 direct helper cases: nonconsecutive turns, all 24 order permutations, late-equalizing ties, boundary clipping, multiple sentences, in-place preservation, NumPy integer labels, empty input and boundary-only contact.
  • 202 passed across tests/test_campplus_utils.py, test_server_app_openai_segments.py, test_realtime_ws_service.py, test_moss_transcribe_diarize_model.py, test_moss_transcribe_diarize_docs.py and test_qwen3_asr_vllm_offline_example.py.
  • Python 3.12.3, existing torch 2.11.0+cu128 environment; no weights loaded, no GPU/model inference or new dependency installation.
  • compileall funasr examples tests and git diff --check passed. Two existing invalid-escape SyntaxWarnings in the unrelated examples/industrial_data_pretraining/fun_asr_nano/tools/format5res.py remain.
  • Independent review found no blocking issue; SSH-signed DCO commit.

Investigated while triaging #3727. This is an independently reproduced assignment defect, not a reproduction or resolution of that reporter's audio/UI symptoms. Keep that issue open for original audio/configuration evidence. No acoustic-accuracy, speaker-clustering or real-world streaming-quality claim.

Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
@LauraGPT
LauraGPT merged commit e1ceba0 into main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant