Skip to content

chore(benchmarks): re-measure the committed datasets and the deck that renders them - #466

Merged
DemchaAV merged 3 commits into
developfrom
chore/refresh-benchmark-datasets
Jul 27, 2026
Merged

chore(benchmarks): re-measure the committed datasets and the deck that renders them#466
DemchaAV merged 3 commits into
developfrom
chore/refresh-benchmark-datasets

Conversation

@DemchaAV

Copy link
Copy Markdown
Owner

Why

Both committed benchmark result files predated the 2.x line.

examples/src/main/resources/benchmarks/comparative.json — the file the flagship deck renders — had one commit in its entire history, taken on the 1.7.1 codebase, before the module split. baselines/current-speed-full.json was captured a week earlier. Meanwhile the deck PDFs built from them are re-committed at every release, most recently by the v2.1.0 release commit, so freshly dated assets carried figures from a codebase two minor versions behind. The deck's own slide printed 2026-06-15 and asked for a refresh before release claims; that never happened across two releases.

What

Re-measured on one quiet machine: five CurrentSpeedBenchmark runs in the full profile aggregated to a median for the baseline (matching how the previous one was captured), and one ComparativeBenchmark run for the deck.

The ratios moved. Unlike absolute timings these compare across machines, because both libraries are measured in the same run on the same hardware:

1000-row fixture published snapshot this run
time advantage vs iText 9 5.38x 4.54x
allocation advantage vs iText 9 4.19x 3.56x
time gap vs JasperReports 1.07x 1.01x

Part of the iText movement is the dependency itself: it went 9.6.0 -> 9.7.0 in 4e148a99, after the snapshot was taken. JasperReports is unchanged at 7.0.7, and the advantage over it narrowed too — most at the small end, 2.02x -> 1.74x at 40 rows.

Two deck footnotes are corrected while the deck is being regenerated. One still asked to "refresh both data and provenance on the final 2.0 commit before release claims" two releases after 2.0 shipped; the other spoke for 2.0 alone. Both now state what the data is: one machine, one run, ratios that travel between machines and timings that do not.

Tests

Full eight-module reactor clean verify — BUILD SUCCESS. examples suite 41/41. The CI performance gate (CurrentSpeedBenchmark in the smoke profile, gate enabled) passes.

The layout snapshot is re-recorded for the reworded footnotes only. The figures alone did not move it: every value kept its digit count, so the auto-width columns kept their widths — verified by running the snapshot test against the fresh data before touching any prose. That is luck rather than robustness; a future refresh that crosses a digit boundary will shift the geometry.

Not addressed here

The baseline still records no machine identity — only profile, iteration counts and a timestamp. That is what makes any comparison against it a measurement of the hardware rather than the code, and it is why this refresh does not, by itself, make the file portable. Worth a follow-up that stamps CPU/OS/JDK/commit into the report writer.

DemchaAV added 2 commits July 26, 2026 23:14
…t renders them

Both committed result files predated the 2.x line. The comparative snapshot the
flagship deck renders had one commit in its entire history, taken on the 1.7.1
codebase, and the perf baseline was captured a week before that — while the deck
PDFs built from them were re-committed at every release, so freshly dated assets
carried figures from a codebase two minor versions behind.

Re-measured on one quiet machine: five current-speed runs aggregated to a median
for the baseline, and one comparative run for the deck.

The ratios moved, and unlike absolute timings they compare across machines
because both libraries are measured in the same run. At 1000 rows the advantage
over iText 9 reads 4.54x on time and 3.56x on allocation, where the published
snapshot said 5.38x and 4.19x; part of that is iText moving 9.6.0 -> 9.7.0 after
the snapshot was taken. Against JasperReports, whose version is unchanged, the
gap at 1000 rows is now 1.1%.

The deck's own footnotes are corrected while they are being regenerated: one
still asked for a refresh "on the final 2.0 commit" two releases after 2.0
shipped, and the summary line still spoke for 2.0 alone. Both now say what the
data is — one machine, one run, ratios that travel and timings that do not.

Layout snapshot re-recorded for the reworded footnotes; the figures alone did
not move it, because every value kept its digit count and the auto-width columns
kept their widths.
The class javadoc carried the same stale instruction the slides did - refresh
data and provenance "on the release commit" for a 2.0 asset, two releases after
2.0 shipped - and it also described the scaling iteration metadata as missing
from a "historical file" rather than as covering only one tier, which is what
the format actually does.

It now says what the slides say: read the ratios rather than the timings,
because both libraries are measured in the same run and their proportions
survive a change of machine while the milliseconds do not.
render-pptx ships two backends: PptxSemanticBackend, the slide-safe skeleton,
and PptxFixedLayoutBackend, which implements FixedLayoutRenderer and is
discovered through FixedLayoutBackendProvider exactly as the PDF backend is.
The module table listed render-pptx as "manifest only" and the architecture
proof card said PPTX was a skeleton, contradicting the deck's own first page,
which already labels render-pptx "fixed layout", and the capability matrix,
which gives PPTX (fixed) a column beside PDF.

The deck is republished by this change, so it was about to carry that
understatement of the engine's own capability into the release assets again.
@DemchaAV
DemchaAV merged commit fbc23cf into develop Jul 27, 2026
10 checks passed
@DemchaAV
DemchaAV deleted the chore/refresh-benchmark-datasets branch July 27, 2026 08:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant