Skip to content

Symmetric Rust and C# benchmarks of Doublets vs SQLite by category, language, address/id space and size - #108

Merged
konard merged 25 commits into
mainfrom
issue-107-bce58c21006b
Oct 5, 2026
Merged

konard merged 25 commits into
mainfrom
issue-107-bce58c21006b

Conversation

@konard

@konard konard commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Fixes #107

What changed

  • One hierarchy in both READMEs: category (links, objects) → language (Rust doublets vs SQLite, C# doublets vs SQLite) → 32/64 bit address/id space benchmarks → size. scripts/benchmark_report.py regenerates the section between the markers from the JSON reports, with one chart per category, language and bits. Groups without results yet show _No results yet._.
  • Symmetrical coverage: Rust (doublets 0.5.0, rusqlite 0.40.2 with bundled SQLite 3.53.2, edition 2024) and C# (net10.0, Platform.Data.Doublets 0.18.1, Platform.Data.Doublets.Sequences 0.6.5, Microsoft.Data.Sqlite 10.0.12) run the same workloads on the same deterministic data with 32 bit (u32/uint) and 64 bit (u64/ulong) ids:
    • links: SQLite links(id INTEGER PRIMARY KEY, "from", "to") with the indices ("from", "to") and ("to", "from"); create, read all, read by id, search by (from, to), read by from, read by to, update and delete;
    • objects: blog posts; create, read all, read by id and delete. Every Doublets store runs with and without a sequences cache (Cached/Uncached);
    • storages: SQLite Memory/File and Doublets United/Split × Volatile/NonVolatile. Each Doublets variant is compared with the SQLite variant of the same durability.
  • No false positives or negatives:
    • every repetition runs on a fresh store in an empty directory;
    • each operation is one timed transaction, and point operations visit records in a scattered order;
    • every result is checked against the expected count and an order-sensitive checksum, so a storage that loses, duplicates or mixes up records fails the run;
    • a difference is reported only if the interquartile ranges don't overlap and the medians differ by more than 5% (otherwise ≈ same).
  • The same work for both engines in both languages:
    • both languages read the inserted id with last_insert_rowid (RETURNING id makes SQLite inserts about 3× slower, see experiments/sqlite_returning);
    • both open transactions with BEGIN IMMEDIATE;
    • both share strings instead of copying them, and checksums don't allocate, so only the stores are timed;
    • warm-up repetitions run for at least a second, so the .NET tiered JIT is finished before timing starts (experiments/csharp_warm_up_order.sh).
  • Library defects found and worked around, each documented in the README with an experiment:
    • Rust split stores lose a link that is updated to reference itself (experiments/split_store_delete);
    • Rust stores break after platform-mem grows their memory, at 1,040,384 links (memory::Whole, experiments/unit_store_growth);
    • C# Delete(id) leaves links in the index trees, so Delete(id, handler: null) is used (experiments/csharp_tree_delete);
    • the default C# size balanced trees degenerate, so the united stores use AVL trees (experiments/csharp_objects_profile);
    • the C# split stores keep their default linked list (experiments/csharp_split_linked_list).
  • Short code, latest dependencies: Rust and C# each have one harness, dataset and storage set. The obsolete BenchmarkDotNet/Criterion projects, nightly toolchain and git dependencies are gone. Dependabot covers cargo, NuGet and actions. The new rust/Cargo.lock no longer has the libsqlite3-sys, atty and remove_dir_all versions Dependabot flags on main.
  • CI:
    • stable Rust (fmt, clippy -D warnings, tests on Linux, macOS and Windows) and .NET 10 (-warnaserror, tests) checks;
    • the Benchmarks workflow measures every table (language, category, bits, size) in its own job, with all of its variants on the same runner;
    • pull requests check the whole pipeline on 1,000 records, and pushes to main publish the README results and charts.
    • sizes: links with 100,000, 1,000,000 and 10,000,000 records, objects with 100,000 and 1,000,000 records. These are the largest that fit into a 6-hour job:
      • in probe run 37228604862, at 100,000,000 links one SQLite Memory repetition took 1–1.4 h, and one SQLite File repetition did not finish in the remaining 4.5 h, so even one job per variant would not fit;
      • 1,000,000 blog posts already take 53–66 minutes in C#, so 10,000,000 would need several jobs per table, and their variants would then be compared across machines.
  • Static analysis: CodeFactor (StyleCop, complexity), Codacy SonarC#, markdownlint and pydocstyle issues are fixed. Codacy enables both D212 and D213, which contradict each other for multi-line docstrings, so the report script uses one-line docstrings.

The README results come from run 37226110636, which ran before the last_insert_rowid, warm-up and shared-strings fixes. The first push to main regenerates them.

Reproduction and regression coverage

  • split_store_keeps_links_updated_to_reference_themselves / SplitStoreKeepsLinksUpdatedToReferenceThemselves: a split store keeps a link updated to (id, id).
  • stores_keep_their_links_when_their_memory_grows and grow_returns_the_whole_memory: Rust stores keep working past 1,040,384 links.
  • every_links_variant_passes_the_lifecycle / EveryLinksVariantPassesTheLifecycle and the objects equivalents run every operation of every variant with checksum validation. In C#, this covers the Delete(id, handler: null) and AVL-tree workarounds.
  • LifecycleRejectsAStorageThatLosesRecords: a storage that loses a record fails instead of looking fast.
  • every_objects_storage_returns_the_stored_posts / EveryObjectsStorageReturnsTheStoredPosts: titles and contents round-trip with and without the cache, including the empty string.
  • Report tests:
    • test_outlier_repetitions_do_not_hide_a_difference uses real C# samples where the old min–max overlap called a 5× difference ≈ same;
    • test_rounding_up_moves_to_the_next_unit_instead_of_an_exponent covers the 999.6 ns → 1e+03 ns bug;
    • test_hierarchy_is_category_language_bits_size checks the README hierarchy.

Verification

  • cargo fmt --check, cargo clippy --all-targets --release -- -D warnings, cargo test --release: 9 tests
  • dotnet build -c Release -warnaserror, dotnet test --project SQLiteVSDoublets.Tests -c Release --no-build: 28 tests
  • python -m unittest discover -s scripts: 17 tests, including chart generation; markdownlint on both READMEs
  • CI on the last commit: Rust (3 OSes), csharp (3 OSes), the Benchmarks pipeline on 1,000 records for all 8 tables, CodeFactor and Codacy all pass
  • full benchmark run 37226110636: all 20 tables

This is benchmark and documentation infrastructure; there are no UI changes or screenshots.

Adding .gitkeep for PR creation (default mode).
This file will be removed when the task is complete.

Issue: #107
@konard konard self-assigned this Oct 4, 2026
konard added 12 commits October 4, 2026 17:05
- One CLI runs links or objects benchmarks for 32 or 64 bit ids at any size
- SQLite stores links(id, from, to) with (from, to) and (to, from) indexes
- Doublets objects mirror Platform.Data.Doublets.Sequences, with and without a sequences cache
- Every timed batch is validated by count and order-sensitive checksum
- Work around a doublets 0.5.0 split store bug with links updated to reference themselves
Replace the legacy BenchmarkDotNet project with a small CLI and an xunit v3 test
project. Both languages now share the dataset, the validated lifecycle, the
variant names and the JSON results format, for 32 and 64 bit ids.

Doublets links are deleted with the Platform.Data.Doublets overload that resets
them first, and objects stores enable external references for raw numbers.
….=10

Raw numbers stored by the objects layout are external references, as in C#.
Repetitions default to 3M / size clamped to 1..=10 in both languages, and
experiments/sizing times one repetition of each variant to plan CI jobs.
The united store's size balanced trees degenerate in the objects workload,
which made every blog post creation linear in the number of stored posts.
experiments/csharp_objects_profile measures it per tree type and
experiments/csharp_tree_delete shows why links are deleted with the resetting
Platform.Data.Doublets overload.
…rarchy

One JSON report per (category, language, bits, size); Doublets are compared
with SQLite of the same durability, and overlapping ranges or medians within
5% are reported as the same.
Repetitions are work / size clamped to 1..=10; objects now use 500,000 posts of
work (5 repetitions at 100,000 posts, 1 at 1,000,000), so that the slowest
C# objects table still fits a 6 hour CI job.
…EADME results; dependabot for cargo and actions
doublets 0.5.0 takes the slice returned by RawMem::grow for the whole
memory, but platform-mem 0.3.0 returns only the grown part, so a fresh
store sees 1,040,384 of its 2^20 links and panics when it creates more
(Rust objects at 1,000,000 posts and links at 10,000,000 failed on CI).
memory::Whole returns the whole memory; experiments/unit_store_growth
reproduces the bug and a test grows unit and split stores past 2^20.
…ords too

Manual runs with 100,000,000 records then measure only the links tables.
@konard

konard commented Oct 4, 2026

Copy link
Copy Markdown
Member Author

🚨 Solution Draft Failed

The automated solution draft encountered an error:

CLAUDE stopped: Authentication expired — re-login required [authentication_failed] — Failed to authenticate: OAuth session expired and could not be refreshed

- Re-authenticate the tool: claude /login  (or set ANTHROPIC_API_KEY for API-key billing)
- Once access is restored, resume with the session ID printed above — no work is lost.

🤖 Models used:

  • Tool: Anthropic Claude Code
  • Requested: opus (claude-opus-5)
  • Thinking level: high (~23999 tokens)
  • Model: Claude Opus 5 (claude-opus-5)

📎 Failure log uploaded as Gist (14731KB)


Now working session is ended, feel free to review and add any feedback on the solution draft.

@konard
konard marked this pull request as ready for review October 4, 2026 19:43
@konard

konard commented Oct 4, 2026

Copy link
Copy Markdown
Member Author

📎 Intermediate working-session log (killed session)

This log file contains the complete execution trace of the AI solution draft process.

📎 Log file uploaded as Gist (9581KB)


Now working session is ended, feel free to review and add any feedback on the solution draft.

@konard

konard commented Oct 4, 2026

Copy link
Copy Markdown
Member Author

⚠️ Container OOM event during a failed work session

container OOM event — a process in the task cgroup was killed earlier; the session-end memory reading was 8.1 GB of 11.7 GB RAM available (30.9% used) at 2026-10-04T19:43:41.936Z

  • Exit code: 1
  • Detected at: 2026-10-04T17:52:25.416Z
  • Working session: d47d4467-5d30-4c64-a61e-787b269289da
  • On-kill policy: resume (--on-session-kill=resume)
Kill diagnostics
  • last session memory reading — 8.1 GB of 11.7 GB RAM available (30.9% used) at 2026-10-04T19:43:41.936Z (phase solve_exit)
  • last session V8 heap reading — 108 MB used of 1.6 GB limit (6.8%) at 2026-10-04T19:43:41.936Z (phase solve_exit)
  • last session disk /: 127.0 GB free of 192.7 GB (34.1% used) at 2026-10-04T19:43:41.936Z (phase solve_exit)
  • container reports State.OOMKilled = true (an OOM event hit the container cgroup)
  • host memory now — 8.1 GB of 11.7 GB RAM available (30.4% used)
  • /proc/pressure/memory: some avg10=0.01 avg60=4.22 avg300=6.55 total=343863199
  • $ --status reports memory exhaustion (cgroup-oom-killer): Docker reported State.OOMKilled=true

The work process survived the container OOM event but later exited with a failure. No replacement session was launched.

📎 The intermediate working-session log was uploaded as a separate comment.

Reported by Hive Mind

@konard

konard commented Oct 4, 2026

Copy link
Copy Markdown
Member Author

We need to recover from https://gist.githubusercontent.com/konard/b7b614bf7feef932de7864cff3a5d710/raw/7324688fe747447f55d7854b33961d59d0ffc64a/tmp-hive-mind-log-upload-OLcnzK-sanitized.log.txt, double check that all CI/CD checks will pass like CodeFactor and so on, and that #107 is fully addressed.

@konard
konard marked this pull request as draft October 5, 2026 00:19
@konard

konard commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

🤖 AI Work Session Started

Starting automated work session at 2026-10-05T00:19:21.566Z

The PR has been converted to draft mode while work is in progress.

This comment marks the beginning of an AI work session. Please wait for the session to finish, and provide your feedback.

Runtime: solve v2.33.11 · tool claude · model opus · task image konard/hive-mind-dind:2.33.11@sha256:999480a8daefeffc7d0ca8ca376df0252eba6a57fa7c45ab2e09b35b692b0288

konard added 8 commits October 5, 2026 00:24
…plit the tree delete experiment into methods
Wrap the prose, put the badges below the title, add the blank lines around
headings, fences and tables, and keep the wide generated tables and repeated
hierarchy headings in a markdownlint-disable block that the report script
writes. The benchmarks workflow lints the committed and regenerated READMEs.
Module and comparison docstrings start on the second line (pydocstyle D213).
Required external flag instead of optional parameters (S2360), a static
property for the warm-up size (S2339), a local for the Unicode sequence marker
(S1450), a parameter name that differs from its method (S3872), and named
swapped link ends instead of reordered arguments (S2234), mirrored in Rust.
The default work is a separate statement like in Rust (S3358).
…, BEGIN IMMEDIATE in Rust

C# inserted with RETURNING id, which makes SQLite inserts about three times
slower than the last_insert_rowid that Rust reads (experiments/sqlite_returning:
2.0 vs 6.0 µs per insert), so C# SQLite create looked 2-3x slower than it is
(14.3 -> 5.1 µs per link in memory). C# now reads sqlite3_last_insert_rowid
through SQLitePCLRaw. Rust transactions start with BEGIN IMMEDIATE like
Microsoft.Data.Sqlite's BeginTransaction().
One warm-up repetition left the .NET tiered JIT unfinished, so whichever C#
variant ran first looked up to 3x slower than its identical twin
(experiments/csharp_warm_up_order.sh). Warm-ups now repeat until a second has
passed, and the reports record warm_up_seconds.
…stores are timed

The Rust dataset and sequences cache copied every title and content String,
while C# shares string references; both now share (Arc<str> in Rust), and the
C# checksum no longer allocates a UTF-8 array per string
(experiments/rust_objects_ab.sh compares the working tree with HEAD).
konard added 3 commits October 5, 2026 01:15
experiments/csharp_split_linked_list compares the default useLinkedList=true
with false: updates take about 1.1 instead of 1.8 µs per link, the other
operations about the same time. doublets 0.5.0 (Rust) has no such list.
Codacy enables both pydocstyle rules, and they contradict each other for every
multi-line docstring, so the details move into comments.
… sizes

Run 37228604862 measured links at 100,000,000: one SQLite Memory repetition
took 1-1.4 hours, and one SQLite File repetition did not finish in the
remaining 4.5 hours before the job limit, so even a job per variant would not
fit. 1,000,000 blog posts take 53-66 minutes in C# (run 37226110636), so ten
millions would need several jobs per table, whose variants would then be
compared across machines.
@konard konard changed the title [WIP] We need clearly divide benchmarks categories, double check we fully cover everything evenly and so on Symmetric Rust and C# benchmarks of Doublets vs SQLite by category, language, address/id space and size Oct 5, 2026
@konard
konard marked this pull request as ready for review October 5, 2026 02:01
@konard

konard commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

Working session summary

I recovered the interrupted session, finished issue #107 and marked PR #108 ready for review. Every check passes on the last commit (e500f13), including CodeFactor and Codacy (0 new issues), and GitHub reports the merge state as clean.

Fixes since the recovered session:

  • C# inserts looked about 3× slower than they are. C# read the new id with INSERT … RETURNING id, while Rust used last_insert_rowid. RETURNING costs about 6.0 µs per insert against 2.0 µs (experiments/sqlite_returning), so both languages now use last_insert_rowid. Rust also now opens transactions with BEGIN IMMEDIATE, as C# already did.
  • C# results depended on run order. One warm-up repetition wasn't enough for the .NET JIT to finish optimizing, so whichever variant ran first looked up to 3× slower than its identical twin. Warm-ups now repeat for at least a second, in both languages.
  • Objects timings included work outside the stores. Rust copied every title and content string, while C# shares them. The C# checksum also allocated memory for every string. Both languages now share strings, and the checksums no longer allocate.
  • C# split stores: the README now notes that they keep a d
Rest of the working session summary (1 KB)
efault internal list that Rust doesn't have. With it, C# updates take about 1.1 µs instead of 1.8 µs; other operations are unaffected.
- **Codacy:** it enables two docstring rules that contradict each other for any multi-line docstring, so the report script now uses one-line docstrings.

**Benchmark sizes:** 10,000,000 links and 1,000,000 blog posts are the largest sizes that fit into a 6-hour job, and both READMEs now explain why.
- **100,000,000 links:** the probe run (37228604862) timed out. One SQLite Memory repetition took 1–1.4 hours. One SQLite File repetition then ran for the remaining ~4.5 hours without finishing, so even one job per variant wouldn't fit.
- **10,000,000 blog posts:** 1,000,000 already takes 53–66 minutes in C#. Ten times that would mean splitting one table across several runners, so its variants would be compared across different machines. I chose not to do that, to avoid false differences.

Local checks pass: Rust format, clippy and 9 tests; C# build and 28 tests; 17 Python tests; markdownlint.

The README results tables still come from run 37226110636, which ran before these fixes. They are regenerated automatically on the first push to `main`; the PR description says so.

The PR's `Cargo.lock` also replaces the old Rust dependency versions (`libsqlite3-sys`, `atty`, `remove_dir_all`) behind the three Dependabot alerts on `main`.

PR: https://github.com/linksplatform/Comparisons.SQLiteVSDoublets/pull/108

This summary was automatically extracted from the AI working session output.

@konard

konard commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

🤖 Solution Draft Log

This log file contains the complete execution trace of the AI solution draft process.

💰 Cost: $11.084839

📊 Context and tokens usage:

Claude Opus 5.5: (4 sub-sessions)

  1. 113.0K / 1M (11%) input tokens, 24.8K / 128K (19%) output tokens
  2. 114.0K / 1M (11%) input tokens, 33.3K / 128K (26%) output tokens
  3. 115.4K / 1M (12%) input tokens, 34.3K / 128K (27%) output tokens
  4. 96.9K / 1M (10%) input tokens, 27.0K / 128K (21%) output tokens

Total: (10.3K new + 594.8K cache writes + 17.7M cache reads) input tokens, 167.1K output tokens, $11.046742 cost

Claude Haiku 4.5:

  • 24.9K / 200K (12%) input tokens, 644 / 64K (1%) output tokens

Total: 24.9K input tokens, 644 output tokens, $0.038097 cost

🤖 Models used:

  • Tool: Anthropic Claude Code
  • Requested: opus (claude-opus-5)
  • Thinking level: high (~23999 tokens)
  • Main model: Claude Opus 5.5 (claude-opus-5-5)
  • Additional models:
    • Claude Haiku 4.5 (claude-haiku-4-5-20251001)

📎 Log file uploaded as Gist (7754KB)


Now working session is ended, feel free to review and add any feedback on the solution draft.

@konard

konard commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

✅ Ready to merge

This pull request is now ready to be merged:

  • All CI checks have passed
  • No merge conflicts
  • No pending changes

Monitored by hive-mind with --auto-restart-until-mergeable flag

@konard
konard merged commit f61d583 into main Oct 5, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

We need clearly divide benchmarks categories, double check we fully cover everything evenly and so on

1 participant