Skip to content
Merged
26 changes: 24 additions & 2 deletions docs/tasks/0103-load-allocations.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
---
status: open
status: done
priority: medium
area: perf
depends: [0011]
touches: [src/template_parser.h#InvokeParser, src/lexer.cpp, src/lexer.h, src/expression_parser.cpp]
touches: [src/template_parser.h#InvokeParser, src/lexer.cpp, src/lexer.h, src/expression_parser.cpp, src/template_parser.cpp, src/value_methods.h]
---
# Loading a tag-heavy template allocates about 115 times per tag

Expand All @@ -29,3 +29,25 @@ no operator. Find out why `GetKeyword` runs so often.

**Done when.** `Load/many_tags` allocations and instructions are measured before and
after with `bench/count.py --baseline`, and the load is at least 2x cheaper.

**Done** in [#377](https://github.com/jinja2cpp/Jinja2Cpp/pull/377). `Load/many_tags` went from 44.15M to 21.99M instructions (-50.2%)
and from 34,249 to 16,559 allocations (-52%); every other `Load/` case got 31-50%
cheaper and no `Render/` case changed. What the profile turned out to say:

- `GetKeyword` ran at every level of the precedence chain for every token it looked at,
about 22,000 times per load. The lexer now classifies each symbol once and stores the
keyword in the `Token`; the parsers read `tok.keyword`. The lookup itself goes by first
character instead of a binary search with a `memcmp` per step.
- The generator and token vector are kept across tags (`LexBuffers`), and tokens are
built in place in the vector.
- The precedence levels did not wrap single operands, contrary to the hypothesis above;
their cost was copies: a `Token` (with its `InternalValue`) copied at every level, the
parse result moved out of each level instead of returned in place, and the postfix and
filter parsers called for every operand. Each level now keeps one result variable, and
`ParseUnaryPlusMinus` calls those parsers only when the next token can start them.
- `StatementsParser` copied the whole `Settings` (ten strings) for every statement tag.
- Smaller: `MarkMacroSpecialNames` returns at once outside macros, attribute subscripts
no longer allocate a constant index node, `IsMethodName` binary-searches, the block end
scan skips plain characters, `endif` moves the statement info instead of copying it.

What is left is in [0109](0109-load-costs-round-2.md).
43 changes: 43 additions & 0 deletions docs/tasks/0109-load-costs-round-2.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
---
status: open
priority: low
area: perf
depends: [0103]
touches: [src/filters.cpp#CreateFilter, src/function_base.h#ParseParams, src/expression_evaluator.cpp#ExpressionFilter, src/expression_parser.cpp#ParseFullExpression, src/lexertk.h]
---
# Load costs after 0103: filter construction, wrappers, the lexer

**Problem.** After 0103 `Load/many_tags` costs 21.99M instructions and 16,559
allocations per load (300 lines, each with an `if`/`else`, four expressions, two filters
and a `set`). The callgrind and gperftools profiles (bench/README.md) split it roughly as:

- allocation itself (`malloc`/`free`/`operator new`/`delete`): about a quarter. Most
allocations are expression nodes, which are needed, but some are not:
- every `ParseFullExpression` allocates a `FullExpressionEvaluator`, also when there is
no inline `if` (about 2,100 per load); `{{ }}` adds an `ExpressionRenderer` on top.
The wrapper carries the render-time `CheckStack`, so dropping it needs a look at
the recursion limits (0098);
- each filter costs three or more allocations and about 2,300 instructions at load:
`CreateFilter` copies `CallParamsInfo` by value through a `std::function`,
`FunctionBase::ParseParams` builds `ArgumentInfo` strings (some longer than the SSO
buffer) and a hash map with buckets for `default(...)`;
- string constants are copied (`InternalValue` copy allocates) from the token into the
`ConstantExpression`, because parsers may backtrack over the token list.
- the lexer: `lexertk::generator::process` plus `Lexer::Preprocess` are about 13%, about
300 instructions per token, including a `std::string` and an `InternalValue` for each
identifier.
- destroying the template (part of the `Load/` loop): about 9%, proportional to the
number of nodes.
- `nonstd::expected` construction and destruction across the precedence chain: about 3%.
- glibc `malloc_consolidate`: about 3%, triggered by the large reallocations of the block
list and the root composition while the fast bins hold the previous template's nodes.
Reserving from a count of the template costs more than it saves on text-heavy
templates (`Load/large_static` got 25% slower in a trial).

**Ideas.** Create filters without copying the call parameters and bind their arguments
from a static table; skip the `FullExpressionEvaluator` when there is no `if`, once the
stack check has another home; keep identifier names as source ranges in tokens and make
the string only where a node needs it.

**Done when.** `Load/many_tags` is measured before and after with
`bench/count.py --baseline` and costs at least 25% fewer instructions than after 0103.
3 changes: 2 additions & 1 deletion docs/tasks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -168,8 +168,9 @@ files go under `touches` (0012 and 0024 each rewrite about a hundred rows).
| [0100](0100-render-hot-path-round-2.md) | Render hot path, round 2 | perf | medium | open |
| [0101](0101-single-pass-generator-lists.md) | Lazy filter results are reusable, Python generators are single-pass | parity | low | done |
| [0102](0102-splitter-byte-scan.md) | The template splitter tries every delimiter at every byte of text | perf | medium | done |
| [0103](0103-load-allocations.md) | Loading a tag-heavy template allocates about 115 times per tag | perf | medium | open |
| [0103](0103-load-allocations.md) | Loading a tag-heavy template allocates about 115 times per tag | perf | medium | done |
| [0104](0104-fixed-costs-per-render.md) | Fixed allocations per render and per macro call | perf | medium | open |
| [0105](0105-include-per-render.md) | `include` and `extends` go through the environment's locked cache on every render | perf | medium | open |
| [0106](0106-percent-format-divergences.md) | `%`-format divergences from Python | parity | low | open |
| [0107](0107-string-filter-divergences.md) | String filter divergences found by the 0061 differential | parity | low | open |
| [0109](0109-load-costs-round-2.md) | Load costs after 0103: filter construction, wrappers, the lexer | perf | low | open |
1 change: 1 addition & 0 deletions src/expression_evaluator.h
Original file line number Diff line number Diff line change
Expand Up @@ -354,6 +354,7 @@ class SubscriptExpression : public Expression
private:
struct Index
{
// Null for an attribute, which is looked up by attrName
ExpressionEvaluatorPtr<Expression> expr;
std::string attrName;
bool isAttr = false;
Expand Down
Loading
Loading