Skip to content

Optimize performance further#29

Open
vladkvit wants to merge 8 commits into
mainfrom
perf2
Open

Optimize performance further#29
vladkvit wants to merge 8 commits into
mainfrom
perf2

Conversation

@vladkvit

Copy link
Copy Markdown
Owner

No description provided.

vladkvit added 8 commits July 10, 2026 22:36
Pass 1 counts mainline SAN tokens per game (tokenizer-only, no SAN
parsing) to compute exact output sizes and per-game CSR offsets.
Pass 2 parses
Two-pass parsing replaces the per-chunk Buffers architecture: pass 1
counts SAN tokens per game to size and offset every output array
exactly; pass 2 parses in parallel over small game tasks with rayon
work stealing, each game writing into its precomputed disjoint slice.
No merge step, no reallocs, no per-thread straggler tail. The GIL is
released during both passes.

Breaking Python API changes:
- ParsedGames now exposes flat global arrays directly (boards,
  from_squares, move_offsets, ..., parsed_move_counts); .chunks,
  num_chunks, PyChunkView and the chunk_multiplier parameter are gone.
- Invalid games keep their pre-counted array range with sentinel-
  filled tails; parsed_move_counts records actual moves. PyGameView
  still exposes only the parsed prefix, matching 4.x semantics.
- Buffers/GameVisitor (src/visitor.rs) deleted; core entry point is
  parse_games_flat, shared by the pyo3 layer and the native bench.

Validation: byte-identical outputs vs 4.0.0 on the lichess 2013-07
corpus (293k games, compare_reference.py), cargo test + test.py green.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant