Benchmarks: 1.1.3 vs 1.2.0

How much faster and leaner 1.2.0 is than the released 1.1.3 on the same inputs, measured with a reproducible benchmark: speed-up per scenario, time per scenario and peak memory, each chart followed by its data table.

1.81× faster on 100 MB text result, string · 1.39× geometric mean over 10 scenarios · 58% less peak memory on 100 MB text result, string (1557 MB to 650 MB) · call-template depth 3,000: runs on 1.2.0, error on 1.1.3.

Contents

Method

  • Machine: AMD Ryzen 9 7950X 16-Core Processor, 32 logical cores, 31 GB RAM, Linux 6.18.40.1-microsoft-standard-WSL2 (linux)
  • Node.js 25.9.0; DOM: jsdom 30.1.1 (library runs and both CLIs), @xmldom/xmldom 0.9.12 (xmldom CLI row)
  • Runs: 2 warm-up + 7 measured per version and scenario, each pair in its own process; a run over 120 s counts as a timeout. Scenarios slower than 10 s per run use 1 warm-up + 3 runs
  • Recorded 2026-09-30; the whole run took 6.2 minutes

1.2.0 is the working tree; 1.1.3 is extracted from its git tag (git archive v1.1.3) into the system temporary directory and installed with npm ci --ignore-scripts. Every version and scenario runs in its own child process, so the versions never share a heap, a JIT or a module cache. Library scenarios parse the inputs once with jsdom and time new XSLTProcessor() + importStylesheet() + transformToString() (or reading transformToStream() to the end), with a garbage collection before each run. CLI scenarios time the whole xslt in.xml t.xsl -o out.html process, start-up included. Peak memory is the process's maximum resident set size (process.resourceUsage().maxRSS), so it includes Node.js, jsdom and the parsed documents. The inputs are generated deterministically by scripts/benchmark/inputs.mjs; nothing is downloaded.

Reproduce:

npm ci --ignore-scripts && npm run build
npm run bench                       # writes scripts/benchmark/results.json
node scripts/benchmark/charts.mjs   # redraws docs/benchmarks/*.svg and the tables below

npm run bench -- --runs 3 --only catalog,sort runs a subset. The charts and tables on this page are generated from scripts/benchmark/results.json.

Speed-up

1.2.0 is faster than 1.1.3 in 10 of 10 compared scenarios; the largest speed-up is 1.81× (100 MB text result, string).
Scenario1.1.3 median1.2.0 medianSpeed-up
100 MB text result, string7.14 s3.94 s1.81×
xsl:number level="any", 8,000153 ms96 ms1.59×
following-sibling::x[1], 8,000168 ms113 ms1.50×
Identity transform, 5 MB2.88 s2.13 s1.35×
apply-templates item[@id], 8,000240 ms178 ms1.35×
Sort 20,000 by two keys548 ms407 ms1.34×
Issue #9 catalogue (3 MB HTML)3.73 s2.83 s1.32×
CLI end to end, jsdom4.86 s3.70 s1.31×
call-template depth 90031 ms24 ms1.30×
Muenchian grouping, 8,000157 ms144 ms1.09×
call-template depth 3,000error62 msn/a
100 MB text result, stream1.2.0 only3.89 sn/a
CLI end to end, XSLT_DOM=xmldom1.2.0 only1.42 sn/a
Standalone binary start-up1.2.0 only28 msn/a

Time per scenario

Median wall time of each scenario for both versions; the slowest is 100 MB text result, string (7.14 s on 1.1.3).
Scenario1.1.3 medianp95minwarm-up + runs1.2.0 medianp95minwarm-up + runs
Issue #9 catalogue (3 MB HTML)3.73 s3.84 s3.67 s2 + 72.83 s2.88 s2.79 s2 + 7
apply-templates item[@id], 8,000240 ms253 ms227 ms2 + 7178 ms181 ms163 ms2 + 7
Muenchian grouping, 8,000157 ms179 ms149 ms2 + 7144 ms157 ms139 ms2 + 7
xsl:number level="any", 8,000153 ms178 ms147 ms2 + 796 ms99 ms90 ms2 + 7
following-sibling::x[1], 8,000168 ms173 ms159 ms2 + 7113 ms115 ms104 ms2 + 7
Sort 20,000 by two keys548 ms577 ms531 ms2 + 7407 ms432 ms401 ms2 + 7
Identity transform, 5 MB2.88 s2.92 s2.78 s2 + 72.13 s2.17 s2.08 s2 + 7
call-template depth 90031 ms33 ms28 ms2 + 724 ms25 ms23 ms2 + 7
call-template depth 3,000error62 ms77 ms60 ms2 + 7
100 MB text result, string7.14 s7.20 s7.02 s2 + 73.94 s3.96 s3.88 s2 + 7
CLI end to end, jsdom4.86 s4.89 s4.82 s2 + 73.70 s3.73 s3.65 s2 + 7
100 MB text result, stream1.2.0 only3.89 s3.95 s3.82 s2 + 7
CLI end to end, XSLT_DOM=xmldom1.2.0 only1.42 s1.43 s1.41 s2 + 7
Standalone binary start-up1.2.0 only28 ms30 ms25 ms2 + 7

Peak memory

Peak resident set size of the process of each scenario; the largest is 100 MB text result, string (1.1.3: 1557 MB, 1.2.0: 650 MB).
Scenario1.1.3 peak RSS1.2.0 peak RSSChange
Issue #9 catalogue (3 MB HTML)922 MB636 MB-31%
apply-templates item[@id], 8,000297 MB277 MB-7%
Muenchian grouping, 8,000267 MB264 MB-1%
xsl:number level="any", 8,000269 MB239 MB-11%
following-sibling::x[1], 8,000279 MB260 MB-7%
Sort 20,000 by two keys326 MB307 MB-6%
Identity transform, 5 MB993 MB658 MB-34%
call-template depth 900181 MB185 MB+2%
call-template depth 3,000error217 MBn/a
100 MB text result, string1557 MB650 MB-58%
CLI end to end, jsdom921 MB627 MB-32%
100 MB text result, stream1.2.0 only550 MBn/a
CLI end to end, XSLT_DOM=xmldom1.2.0 only453 MBn/a
Standalone binary start-up1.2.0 onlynot measuredn/a

Scenarios

ScenarioWhat it does
Issue #9 catalogue10,500 courses grouped by category with keys (Muenchian), sorted by title, labels looked up by key from localised strings, format-number(), (a) or (b) tests as in #9; 3 MB of HTML
apply-templates8,000 items, every second one matched by match="item[@id]", the others by match="item"
Muenchian grouping8,000 items in 200 groups: generate-id() = generate-id(key(...)[1]), count() and sum() over each group
xsl:number level="any"8,000 elements numbered with count="i[@k='1']"
following-sibling::x[1]A loop over 8,000 siblings reading the next one
Sort20,000 items by a text key, then a numeric key descending
Identity transform@*|node() copied recursively through a 5 MB document
call-template depthRecursive named template, 900 and 3,000 nested template instantiations (the root template included)
100 MB text result100 million characters of method="text" output, returned as one string, and in 1.2.0 also read from transformToStream()
CLI end to endxslt in.xml t.xsl -o out.html on the catalogue, with jsdom; 1.2.0 also with XSLT_DOM=xmldom
Standalone binary start-upxslt --version with the host's single executable (1.2.0 only; built by make binaries when missing). Its memory is not measured: the executable ignores the --import hook used to read it

Notes and caveats

  • jsdom dominates wall time. Every DOM access of the engine crosses jsdom's wrappers; parsing the inputs is excluded from library runs but not from CLI runs. With XSLT_DOM=xmldom the same CLI transformation is much faster, see the CLI rows.
  • Numbers vary by machine. Compare factors, not absolute times, and rerun npm run bench on your own hardware; run-to-run spread is visible in the p95 and min columns.
  • What changed in 1.2.0 (see the CHANGELOG):
    • Serialization walks the result tree on an explicit stack and writes bounded chunks that are joined once (task 0024 and the chunked serializer of task 0010): the identity transform, the catalogue and the 100 MB text result are faster, and the 100 MB result needs well under half the peak memory of 1.1.3.
    • transformToStream() (task 0010) serializes on demand, so the consumer never holds the whole result as one string: the streamed 100 MB row peaks lower than the string row of the same version.
    • Templates run from an explicit work stack (task 0005): 3,000 nested templates work, where 1.1.3 stops at about 1,000 to 1,400 levels with Template recursion too deep.
    • Document order is computed once per transformation (DocumentOrderIndex, task 0007) instead of calling compareDocumentPosition for each comparison.
    • The CLI can use @xmldom/xmldom (task 0007), which starts faster and parses faster than jsdom. jsdom stays the default because it expands entities declared in an internal DTD subset and xmldom does not.
    • Sorting nodes in document order no longer reads every attribute's index in its element, and name tests read the document's content type only when it matters. Before that fix Muenchian grouping was the one scenario slower than 1.1.3 (about 0.88×); it is now faster too.
  • Most of the large algorithmic fixes (linear template matching, keys, xsl:number, axis::x[n]) already shipped in 1.1.2 and 1.1.3, so both versions here are within a small factor of each other on those scenarios; the gains of 1.2.0 are in serialization, memory, recursion depth and the CLI.

This page is generated from docs/BENCHMARKS.md in the repository. Corrections are welcome as a pull request to that file.