Rewrite semantics¶
{ #architecture-rewrite-semantics }
Stable ID: ARCHITECTURE-REWRITE-SEMANTICS
Purpose¶
Document design decisions that determine how the rewrite semantics are implemented. For the domain model of changes and the rewrite step, see Rewrite semantics.
Units versus text ranges¶
The rules that combine changes can be stated over the text ranges that the changes affect, or over the consecutive units - characters or AST nodes - that the changes act on. Text ranges look like the simpler notion, as every change ultimately affects a range of text.
A pure text-range formulation, however, fails for the dominance rule. That rule ignores every change whose text range is contained in the text range of a replacement, yet a prepend or an append on exactly the units that are replaced must still be applied, even though its location lies within the text range of that replacement. Stated over text ranges, the rule thus needs an exception that cannot be expressed in text ranges alone.
Stated over units the exception disappears: a change is ignored when the units of a replacement contain its units, and insertions on the same units are not contained, hence not ignored. The AST-based view gives a second reason: a node and one of its descendants can have equal text ranges - e.g., a declaration statement node and the declaration node it contains in the CDT parser - so their text ranges cannot tell them apart, whereas their units can.
Decision:
- The combination rules are stated over the consecutive units that the changes act on.
- Text ranges are used to describe the consequences of those rules, never to define them.
- For zero consecutive units, which contain no units at all, the location determines identity.
Impossible combinations of changes¶
The collected changes might contain changes whose combination is not possible. For example, it is impossible to replace the same AST node with different texts. The following behaviours are possible when such a situation occurs:
- keep-first: keep the first collected change and ignore later ones
- keep-last: drop earlier collected changes and keep the last one
- reject: reject the combination and stop processing.
Diagnostics may be produced in addition to the selected behaviour (e.g., warnings for keep-first/keep-last, errors for reject).
The reject behaviour is independent of the order in which changes are collected, whereas keep-first and keep-last depend on that order.
Decision:
For simplicity, we have decided that
- Impossible changes are rejected.
Note that removing two overlapping sequences of nodes is treated the same as replacing them and thus considered an impossible change.
Combining changes on nodes from different roots is also an impossible change - see Position consistency for why this is rejected rather than applied.
Repeated changes¶
In the collected changes, the same changes might occur more than once. To give some examples,
- a node is removed multiple times, and
- a node is prepended with the same text multiple times.
The behaviour of a change could be idempotent or not. Furthermore, who makes this decision? The user or the framework?
Decision: Based on our experience with Renaissance that many features were never needed by high-quality transformations, we have decided to keep the rules as simple as possible.
- The behaviour of changes is not configurable by the user, but fixed by the framework.
- None of the changes is idempotent.
Note that non-idempotence results in quite different behaviours, depending on the kind of change:
-
Replacements: replacing the same node twice is considered an impossible combination - even when the replacement is the same.
-
Insertions: The same text can be appended multiple times to the same node. The text will be inserted as many times as the number of appends.
Combinations at the same text location¶
Choices
- Implementation freedom
- Specify
We have chosen to specify as it makes the outcome more predictable and repeatable.
How to specify order? Options include
- (reversed) order of insertion,
- (reversed) alphabetically,
- AST-based
We have chosen
- When different AST-nodes are involved to order based on AST structure, e.g., for prepends, appends, and surrounds of different AST nodes
- When the same AST-node is involved to order based on the (reversed) order of insertion into the collection of changes - see Particular combinations in the concept page for how the direction depends on the operator (prepends/surround-before-texts vs. appends/surround-after-texts) and for concrete examples.
TO BE Removed?¶
The following text is already covered by previous remarks, yet without some details. Are the details relevant, useful for our developers?
Corner case: Identical replacements¶
Replacing the same AST node with different texts is not possible. For the corner case of replacing the same AST node more than once with the same text, different behaviours are possible:
- error: the combination of replacements is considered invalid and an error is raised.
- idempotent: the combination of replacements is considered valid and the node is replaced by the text.
- warning: the combination of replacements is considered suspicious, a warning is raised while the node is replaced by the text.
Decision:
This framework adopts the error behaviour for replacements. In other words, replacing the same AST node more than once is considered an error, regardless of whether the replacement text is identical. Note that as this framework considers removing an AST node equal to replacing that AST node with an empty string, removing the same AST node more than once is considered an error.
Corner case: Identical insertions¶
Prepending, appending, and surrounding different texts before, around, and after an AST node is possible.
For the corner case of using the same insertion operator on the same AST node with the same text more than once, different behaviours are possible:
- idempotent: the text is inserted only once.
- non-idempotent: the text is inserted more than once - as often as the insertion operator is used.
Decision:
This framework adopts non-idempotent insertion semantics: applying append, prepend, or surround multiple times to the same AST node results in repeated insertions, even when the inserted text is identical.