The transcribe cli seems to be crashing on me and I wonder what could be wrong. This works great on this file when using whisper models, but then I don't get diarization. Please let me know how I can help you troubleshoot or if you need more info.
Here are the file sizes. I downloaded these from the links listed within this repo:
1833665696 MOSS-Transcribe-Diarize-F16.gguf
986899616 MOSS-Transcribe-Diarize-Q8_0.gguf
[info] ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices
[info] ggml_metal_library_init: using embedded metal library
[info] ggml_metal_library_init: loaded in 0.007 sec
[info] ggml_metal_rsets_init: creating a residency set collection (keep_alive = 180 s)
[info] ggml_metal_device_init: GPU name: MTL0 (Apple M4)
[info] ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
[info] ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
[info] ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
[info] ggml_metal_device_init: simdgroup reduction = true
[info] ggml_metal_device_init: simdgroup matrix mul. = true
[info] ggml_metal_device_init: has unified memory = true
[info] ggml_metal_device_init: has bfloat = true
[info] ggml_metal_device_init: has tensor = false
[info] ggml_metal_device_init: use residency sets = true
[info] ggml_metal_device_init: use shared buffers = true
[info] ggml_metal_device_init: recommendedMaxWorkingSetSize = 26800.60 MB
[info] ggml_metal_init: allocating
[info] ggml_metal_init: found device: Apple M4
[info] ggml_metal_init: picking default device: Apple M4
[info] ggml_metal_init: use fusion = true
[info] ggml_metal_init: use concurrency = true
[info] ggml_metal_init: use graph optimize = true
[info] moss: using metal backend: MTL0
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_im2col_f32', name = 'kernel_im2col_f32'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_im2col_f32 0x105305750 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_mul_mm_f32_f32', name = 'kernel_mul_mm_f32_f32_bci=1_bco=1_ne12=1_ne13=1_r2=1_r3=1'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_mul_mm_f32_f32_bci=1_bco=1_ne12=1_ne13=1_r2=1_r3=1 0x105305fd0 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_bin_fuse_f32_f32_f32', name = 'kernel_bin_fuse_f32_f32_f32_op=0_nf=1_rb=0_cb=1'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_bin_fuse_f32_f32_f32_op=0_nf=1_rb=0_cb=1 0x105306850 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_unary_f32_f32_4', name = 'kernel_unary_f32_f32_4_op=104_cnt=0'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_unary_f32_f32_4_op=104_cnt=0 0x1053070d0 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_mul_mm_f32_f32', name = 'kernel_mul_mm_f32_f32_bci=0_bco=1_ne12=1_ne13=1_r2=1_r3=1'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_mul_mm_f32_f32_bci=0_bco=1_ne12=1_ne13=1_r2=1_r3=1 0x105307950 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_cpy_f32_f32', name = 'kernel_cpy_f32_f32'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_cpy_f32_f32 0x105307cd0 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_bin_fuse_f32_f32_f32_4', name = 'kernel_bin_fuse_f32_f32_f32_4_op=0_nf=1_rb=0_cb=0'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_bin_fuse_f32_f32_f32_4_op=0_nf=1_rb=0_cb=0 0x105308050 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_norm_mul_add_f32_4', name = 'kernel_norm_mul_add_f32_4'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_norm_mul_add_f32_4 0x1053083d0 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_mul_mm_f16_f32', name = 'kernel_mul_mm_f16_f32_bci=0_bco=1_ne12=1_ne13=1_r2=1_r3=1'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_mul_mm_f16_f32_bci=0_bco=1_ne12=1_ne13=1_r2=1_r3=1 0xbed04c000 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_flash_attn_ext_pad', name = 'kernel_flash_attn_ext_pad_mask=0_ncpsg=64'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_flash_attn_ext_pad_mask=0_ncpsg=64 0xbed04c380 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_flash_attn_ext_f32_dk64_dv64', name = 'kernel_flash_attn_ext_f32_dk64_dv64_mask=0_sinks=0_bias=0_scap=0_kvpad=1_bcm=0_ns10=64_ns20=64_nsg=4'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_flash_attn_ext_f32_dk64_dv64_mask=0_sinks=0_bias=0_scap=0_kvpad=1_bcm=0_ns10=64_ns20=64_nsg=4 0xbed04c700 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_unary_f32_f32_4', name = 'kernel_unary_f32_f32_4_op=106_cnt=0'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_unary_f32_f32_4_op=106_cnt=0 0xbed04d180 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_get_rows_f16', name = 'kernel_get_rows_f16'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_get_rows_f16 0xbed04ce00 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_bin_fuse_f32_f32_f32', name = 'kernel_bin_fuse_f32_f32_f32_op=2_nf=1_rb=0_cb=1'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_bin_fuse_f32_f32_f32_op=2_nf=1_rb=0_cb=1 0xbed04d880 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_rms_norm_mul_f32_4', name = 'kernel_rms_norm_mul_f32_4'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_rms_norm_mul_f32_4 0xbed04dc00 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_cpy_f32_f16', name = 'kernel_cpy_f32_f16'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_cpy_f32_f16 0xbed04df80 | th_max = 1024 | th_width = 32
[debug] ggml_metal_library_compile_pipeline: compiling pipeline: base = 'kernel_rope_neox_f32', name = 'kernel_rope_neox_f32_imrope=0_is_back=0'
[debug] ggml_metal_library_compile_pipeline: loaded kernel_rope_neox_f32_imrope=0_is_back=0 0xbed04e300 | th_max = 1024 | th_width = 32
audio: file.wav
samples: 112562167
duration: 7035.135 s
sample rate 16000 Hz mono float32
model: MOSS-Transcribe-Diarize-F16.gguf -> ok
backend: MTL0
name: MOSS-Transcribe-Diarize 0.9B
license: apache-2.0
max audio: 10458.4 s (context 131072 tok, ~14336 MiB KV max)
handy-computer/transcribe.cpp/ggml/src/ggml-metal/ggml-metal-ops.cpp:2696: GGML_ASSERT(ne01 < 65536) failed
(lldb) process attach --pid 44439
Process 44439 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP
frame #0: 0x000000018a005478 libsystem_kernel.dylib`__wait4 + 8
libsystem_kernel.dylib`__wait4:
-> 0x18a005478 <+8>: b.lo 0x18a005498 ; <+40>
0x18a00547c <+12>: pacibsp
0x18a005480 <+16>: stp x29, x30, [sp, #-0x10]!
0x18a005484 <+20>: mov x29, sp
Target 0: (transcribe-cli) stopped.
Executable binary set to "handy-computer/transcribe.cpp/build/bin/transcribe-cli".
Architecture set to: arm64-apple-macosx-.
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP
* frame #0: 0x000000018a005478 libsystem_kernel.dylib`__wait4 + 8
frame #1: 0x00000001050abd28 transcribe-cli`ggml_abort + 156
frame #2: 0x000000010503fbb0 transcribe-cli`ggml_metal_op_flash_attn_ext + 4116
frame #3: 0x0000000105038ef4 transcribe-cli`ggml_metal_op_encode + 1720
frame #4: 0x0000000105038644 transcribe-cli`__ggml_metal_set_n_cb_block_invoke + 188
frame #5: 0x0000000105038200 transcribe-cli`ggml_metal_graph_compute + 588
frame #6: 0x000000010505fe20 transcribe-cli`ggml_backend_sched_graph_compute_async + 2508
frame #7: 0x000000010505f3a0 transcribe-cli`ggml_backend_sched_graph_compute + 28
frame #8: 0x0000000104eef728 transcribe-cli`transcribe::moss::(anonymous namespace)::run(transcribe_session*, float const*, int, transcribe_run_params const*) + 1772
frame #9: 0x0000000104e8dbdc transcribe-cli`run_one_inner(transcribe_session*, float const*, int, transcribe_run_params const*) + 464
frame #10: 0x0000000104e8b634 transcribe-cli`transcribe_run + 116
frame #11: 0x0000000104e86224 transcribe-cli`main + 13796
frame #12: 0x0000000189c83e00 dyld`start + 6992
(lldb) quit
The transcribe cli seems to be crashing on me and I wonder what could be wrong. This works great on this file when using whisper models, but then I don't get diarization. Please let me know how I can help you troubleshoot or if you need more info.
I am using this command:
I also get the same result when using
R="MOSS-Transcribe-Diarize-Q8_0.gguf"Here are the file sizes. I downloaded these from the links listed within this repo:
Error message: