bjoernager/rust - mandelbrot.dk

Author	SHA1	Message	Date
Nicholas Nethercote	0bae33fcd5	Avoid nested replacement ranges. In a case like this: ``` mod a { mod b { #[cfg_attr(unix, inline)] fn f() { #[cfg_attr(linux, inline)] fn g1() {} #[cfg_attr(linux, inline)] fn g2() {} } } } ``` We currently end up with the following replacement ranges. - The lazy tokens for `f` has replacement ranges for `g1` and `g2`. - The lazy tokens for `a` has replacement ranges for `f`, `g1`, and `g2`. I.e. the replacement ranges for `g1` and `g2` are duplicated. In general, replacement ranges for inner AST nodes are duplicated up the chain for each nested `collect_tokens` call. And the code that processes the replacements is careful about the ordering in which the replacements are applied, to ensure that inner replacements are applied before outer replacements. But all of this is unnecessary. If you apply an inner replacement and then an outer replacement, the outer replacement completely overwrites the inner replacement. This commit avoids the duplication by removing replacements from `self.capture_state.parser_replacements` when they are used. (The effect on the example above is that the lazy tokesn for `a` no longer include replacement ranges for `g1` and `g2`.) This eliminates the possibility of nested replacements on individual AST nodes, which avoids the need for careful ordering of replacements.	2024-08-23 14:40:08 +10:00
Nicholas Nethercote	1ae521e9d5	Return earlier in some cases in `collect_token`. This example triggers an assertion failure: ``` fn f() -> u32 { #[cfg_eval] #[cfg(not(FALSE))] 0 } ``` The sequence of events: - `configure_annotatable` calls `parse_expr_force_collect`, which calls `collect_tokens`. - Within that, we end up in `parse_expr_dot_or_call`, which again calls `collect_tokens`. - The return value of the `f` call is the expression `0`. - This inner call collects tokens for `0` (parser range 10..11) and creates a replacement covering `#[cfg(not(FALSE))] 0` (parser range 0..11). - We return to the outer `collect_tokens` call. The return value of the `f` call is again the expression `0`, again with the range 10..11, but the replacement from earlier covers the range 0..11. The code mistakenly assumes that any attributes from an inner `collect_tokens` call fit entirely within the body of the result of an outer `collect_tokens` call. So it adjusts the replacement parser range 0..11 to a node range by subtracting 10, resulting in -10..1. This is an invalid range and triggers an assertion failure. It's tricky to follow, but basically things get complicated when an AST node is returned from an inner `collect_tokens` call and then returned again from an outer `collect_token` node without being wrapped in any kind of additional layer. This commit changes `collect_tokens` to return early in some extra cases, avoiding the construction of lazy tokens. In the example above, the outer `collect_tokens` returns earlier because the `0` token already has tokens and `self.capture_state.capturing` is `Capturing::No`. This early return avoids the creation of the invalid range and the assertion failure. Fixes #129166. Note: these invalid ranges have been happening for a long time. #128725 looks like it's at fault only because it introduced the assertion that catches the invalid ranges.	2024-08-23 14:40:08 +10:00
Nicholas Nethercote	312ecdb2ed	Avoid unnecessary `cloned`.	2024-08-23 14:40:08 +10:00
Nicholas Nethercote	deab741ab4	Clarify a comment.	2024-08-23 14:40:08 +10:00
Nicholas Nethercote	9d31f86f0d	Overhaul token collection. This commit does the following. - Renames `collect_tokens_trailing_token` as `collect_tokens`, because (a) it's annoying long, and (b) the `_trailing_token` bit is less accurate now that its types have changed. - In `collect_tokens`, adds a `Option<CollectPos>` argument and a `UsePreAttrPos` in the return type of `f`. These are used in `parse_expr_force_collect` (for vanilla expressions) and in `parse_stmt_without_recovery` (for two different cases of expression statements). Together these ensure are enough to fix all the problems with token collection and assoc expressions. The changes to the `stringify.rs` test demonstrate some of these. - Adds a new test. The code in this test was causing an assertion failure prior to this commit, due to an invalid `NodeRange`. The extra complexity is annoying, but necessary to fix the existing problems.	2024-08-16 09:07:55 +10:00
Nicholas Nethercote	c8098be41f	Convert a bool to `Trailing`. This pre-existing type is suitable for use with the return value of the `f` parameter in `collect_tokens_trailing_token`. The more descriptive name will be useful because the next commit will add another boolean value to the return value of `f`.	2024-08-16 09:07:29 +10:00
Nicholas Nethercote	55906aa240	Make visibilities minimal and consistent in `attr_wrapper.rs`.	2024-08-16 09:06:15 +10:00
Nicholas Nethercote	af0093a6b8	Remove size assertion on `AttrWrapper`. It's not an important type when it comes to memory use.	2024-08-16 09:06:15 +10:00
Nicholas Nethercote	d1f05fd184	Distinguish the two kinds of token range. When collecting tokens there are two kinds of range: - a range relative to the parser's full token stream (which we get when we are parsing); - a range relative to a single AST node's token stream (which we use within `LazyAttrTokenStreamImpl` when replacing tokens). These are currently both represented with `Range<u32>` and it's easy to mix them up -- until now I hadn't properly understood the difference. This commit introduces `ParserRange` and `NodeRange` to distinguish them. This also requires splitting `ReplaceRange` in two, giving the new types `ParserReplacement` and `NodeReplacement`. (These latter two names reduce the overloading of the word "range".) The commit also rewrites some comments to be clearer. The end result is a little more verbose, but much clearer.	2024-08-01 19:30:40 +10:00
Nicholas Nethercote	2eb2ef1684	Streamline attribute stitching on AST nodes. It can be done more concisely.	2024-08-01 19:30:32 +10:00
Nicholas Nethercote	84ac80f192	Reformat `use` declarations. The previous commit updated `rustfmt.toml` appropriately. This commit is the outcome of running `x fmt --all` with the new formatting options.	2024-07-29 08:26:52 +10:00
Trevor Gross	af52be2cea	Rollup merge of #128224 - nnethercote:fewer-replace_ranges, r=petrochenkov Remove unnecessary range replacements This PR removes an unnecessary range replacement in `collect_tokens_trailing_token`, and does a couple of other small cleanups. r? ````@petrochenkov````	2024-07-26 19:03:06 -04:00
Nicholas Nethercote	55d37ae711	Remove an unnecessary block.	2024-07-26 17:37:03 +10:00
Nicholas Nethercote	6ea2da5a28	Tweak a loop. A fully imperative style is easier to read than a half-iterator, half-imperative style. Also, rename `inner_attr` as `attr` because it might be an outer attribute.	2024-07-26 17:37:03 +10:00
Nicholas Nethercote	6e87858f26	Fix a comment. Imagine you have replace ranges (2..20,X) and (5..15,Y), and these tokens: ``` a,b,c,d,e,f,g,h,i,j,k,l,m,n,o,p,q,r,s,t,u,v,w,x ``` If we replace (5..15,Y) first, then (2..20,X) we get this sequence ``` a,b,c,d,e,Y,_,_,_,_,_,_,_,_,_,p,q,r,s,t,u,v,w,x a,b,X,_,_,_,_,_,_,_,_,_,_,_,_,_,_,_,_,_,u,v,w,x ``` which is what we want. If we do it in the other order, we get this: ``` a,b,X,_,_,_,_,_,_,_,_,_,_,_,_,p,q,r,s,t,u,v,w,x a,b,X,_,_,Y,_,_,_,_,_,_,_,_,_,_,_,_,_,_,u,v,w,x ``` which is wrong. So it's true that we need the `.rev()` but the comment is wrong about why.	2024-07-26 17:37:03 +10:00
Nicholas Nethercote	a560810a69	Don't include inner attribute ranges in `CaptureState`. The current code is this: ``` self.capture_state.replace_ranges.push((start_pos..end_pos, Some(target))); self.capture_state.replace_ranges.extend(inner_attr_replace_ranges); ``` What's not obvious is that every range in `inner_attr_replace_ranges` must be a strict sub-range of `start_pos..end_pos`. Which means, in `LazyAttrTokenStreamImpl::to_attr_token_stream`, they will be done first, and then the `start_pos..end_pos` replacement will just overwrite them. So they aren't needed.	2024-07-26 14:18:20 +10:00
Nicholas Nethercote	e631b1ebfa	Invert the sense of `is_complete` and rename it as `needs_tokens`. I have always found `is_complete` an unhelpful name. The new name (and inverted sense) fits in better with the conditions at its call sites.	2024-07-26 09:58:34 +10:00
Nicholas Nethercote	3d363c3d99	Move `is_complete` to the module that uses it. And make it non-`pub`.	2024-07-26 09:44:39 +10:00
Nicholas Nethercote	4288edb219	Inline and remove `AttrWrapper::is_complete`. It has a single call site. This change makes the two `needs_collect` conditions more similar to each other, and therefore easier to understand.	2024-07-26 09:44:07 +10:00
Nicholas Nethercote	caee195bdd	Invert early exit conditions in `collect_tokens_trailing_token`. This has been bugging me for a while. I find complex "if any of these are true" conditions easier to think about than complex "if all of these are true" conditions, because you can stop as soon as one is true.	2024-07-26 09:43:41 +10:00
Nicholas Nethercote	1dd566a6d0	Overhaul comments in `collect_tokens_trailing_token`. Adding details, clarifying lots of little things, etc. In particular, the commit adds details of an example. I find this very helpful, because it's taken me a long time to understand how this code works.	2024-07-19 15:25:55 +10:00
Nicholas Nethercote	f9c7ca70cb	Move `inner_attr` code downwards. This puts it just before the `replace_ranges` initialization, which makes sense because the two variables are closely related.	2024-07-19 15:25:54 +10:00
Nicholas Nethercote	1f67cf9e63	Remove `final_attrs` local variable. It's no shorter than `ret.attrs()`, and `ret.attrs()` is used multiple times earlier in the function.	2024-07-19 15:25:54 +10:00
Nicholas Nethercote	757f73f506	Simplify `CaptureState::inner_attr_ranges`. The `Option`s within the `ReplaceRange`s within the hashmap are always `None`. This PR omits them and inserts them when they are extracted from the hashmap.	2024-07-19 15:25:54 +10:00
Nicholas Nethercote	487802d6c8	Remove `TrailingToken`. It's used in `Parser::collect_tokens_trailing_token` to decide whether to capture a trailing token. But the callers actually know whether to capture a trailing token, so it's simpler for them to just pass in a bool. Also, the `TrailingToken::Gt` case was weird, because it didn't result in a trailing token being captured. It could have been subsumed by the `TrailingToken::MaybeComma` case, and it effectively is in the new code.	2024-07-18 17:28:49 +10:00
Jubilee	125343e7ab	Rollup merge of #127558 - nnethercote:more-Attribute-cleanups, r=petrochenkov More attribute cleanups A follow-up to #127308. r? ```@petrochenkov```	2024-07-13 20:19:46 -07:00
Nicholas Nethercote	8a390bae06	Change empty replace range condition. The new condition is equivalent in practice, but it's much more obvious that it would result in an empty range, because the condition lines up with the contents of the iterator.	2024-07-10 14:41:39 +10:00
Nicholas Nethercote	f5527949f2	Move `Spacing` into `FlatToken`. It's only needed for the `FlatToken::Token` variant. This makes things a little more concise.	2024-07-09 21:54:32 +10:00
Nicholas Nethercote	a88c4d67d9	Split the stack in `make_attr_token_stream`. It makes for shorter code, and fewer allocations.	2024-07-08 19:04:13 +10:00
Nicholas Nethercote	b16201317e	Use iterator normally in `make_attr_token_stream`. In a `for` loop, instead of a `while` loop.	2024-07-08 19:04:13 +10:00
Nicholas Nethercote	a47ae57a18	Use an `@` pattern to shorten some code.	2024-07-08 19:03:50 +10:00
Nicholas Nethercote	99721c8469	Clear `inner_attr_ranges` regularly. There's a comment saying we don't do it for performance reasons, but it doesn't actually affect performance. The commit also tweaks the control flow, to make clearer that two code paths are mutually exclusive.	2024-07-08 16:53:10 +10:00
Nicholas Nethercote	022582ca46	Remove `Clone` derive from `LazyAttrTokenStreamImpl`.	2024-07-07 16:24:51 +10:00
Nicholas Nethercote	3a5c4b6e4e	Rename some attribute types for consistency. - `AttributesData` -> `AttrsTarget` - `AttrTokenTree::Attributes` -> `AttrTokenTree::AttrsTarget` - `FlatToken::AttrTarget` -> `FlatToken::AttrsTarget`	2024-07-07 16:14:30 +10:00
Nicholas Nethercote	9d33a8fe51	Simplify `ReplaceRange`. Currently the second element is a `Vec<(FlatToken, Spacing)>`. But the vector always has zero or one elements, and the `FlatToken` is always `FlatToken::AttrTarget` (which contains an `AttributesData`), and the spacing is always `Alone`. So we can simplify it to `Option<AttributesData>`. An assertion in `to_attr_token_stream` can can also be removed, because `new_tokens.len()` was always 0 or 1, which means than `range.len()` is always greater than or equal to it, because `range.is_empty()` is always false (as per the earlier assertion).	2024-07-07 15:58:36 +10:00
Nicholas Nethercote	dd790ab8ef	Remove some unnecessary integer conversions. These should have been removed in #127233 when the positions were changed from `usize` to `u32`.	2024-07-05 08:27:17 +10:00
Nicholas Nethercote	edeebe675b	Import `std::{iter,mem}`.	2024-07-02 20:29:01 +10:00
Nicholas Nethercote	6f6015679f	Rename `make_token_stream`. And update the comment. Clearly the return type of this function was changed at some point in the past, but its name and comment weren't updated to match.	2024-07-02 17:38:43 +10:00
Nicholas Nethercote	3d750e2702	Shrink parser positions from `usize` to `u32`. The number of source code bytes can't exceed a `u32`'s range, so a token position also can't. This reduces the size of `Parser` and `LazyAttrTokenStreamImpl` by eight bytes each.	2024-07-02 17:03:53 +10:00
Nicholas Nethercote	f5b28968db	Move more things around in `collect_tokens_trailing_token`. To make things a little clearer, and to avoid some `mut` variables.	2024-07-02 10:46:44 +10:00
Nicholas Nethercote	8b5a7eb7f4	Move things around in `collect_tokens_trailing_token`. So that the `capturing` state is adjusted immediately before and after the call to `f`.	2024-07-02 10:46:44 +10:00
Nicholas Nethercote	2342770f49	Flip an if/else in `AttrTokenStream::to_attr_token_stream`. To put the simple case first.	2024-07-02 10:46:44 +10:00
Nicholas Nethercote	36c30a968b	Fix comment. Both the indenting, and the missing `)`.	2024-07-02 10:46:44 +10:00
Nicholas Nethercote	d6c0b8117e	Fix a typo in a comment.	2024-07-02 10:46:43 +10:00
bors	894f7a4ba6	Auto merge of #126678 - nnethercote:fix-duplicated-attrs-on-nt-expr, r=petrochenkov Fix duplicated attributes on nonterminal expressions This PR fixes a long-standing bug (#86055) whereby expression attributes can be duplicated when expanded through declarative macros. First, consider how items are parsed in declarative macros: ``` Items: - parse_nonterminal - parse_item(ForceCollect::Yes) - parse_item_ - attrs = parse_outer_attributes - parse_item_common(attrs) - maybe_whole! - collect_tokens_trailing_token ``` The important thing is that the parsing of outer attributes is outside token collection, so the item's tokens don't include the attributes. This is how it's supposed to be. Now consider how expression are parsed in declarative macros: ``` Exprs: - parse_nonterminal - parse_expr_force_collect - collect_tokens_no_attrs - collect_tokens_trailing_token - parse_expr - parse_expr_res(None) - parse_expr_assoc_with - parse_expr_prefix - parse_or_use_outer_attributes - parse_expr_dot_or_call ``` The important thing is that the parsing of outer attributes is inside token collection, so the the expr's tokens do include the attributes, i.e. in `AttributesData::tokens`. This PR fixes the bug by rearranging expression parsing to that outer attribute parsing happens outside of token collection. This requires a number of small refactorings because expression parsing is somewhat complicated. While doing so the PR makes the code a bit cleaner and simpler, by eliminating `parse_or_use_outer_attributes` and `Option<AttrWrapper>` arguments (in favour of the simpler `parse_outer_attributes` and `AttrWrapper` arguments), and simplifying `LhsExpr`. r? `@petrochenkov`	2024-06-19 13:58:21 +00:00
Nicholas Nethercote	1fbb3eca67	Expand another comment.	2024-06-19 18:53:24 +10:00
Oli Scherer	c91edc3888	Prefer `dcx` methods over fields or fields' methods	2024-06-18 13:45:08 +00:00
Nicholas Nethercote	0d97669a17	Simplify `static_assert_size`s. We want to run them on all 64-bit platforms.	2024-04-18 15:36:25 +10:00
Zalathar	2d47cd77ac	Check `x86_64` size assertions on `aarch64`, too This makes it easier for contributors on aarch64 workstations (e.g. Macs) to notice when these assertions have been violated.	2024-04-03 16:53:03 +11:00
Nicholas Nethercote	80d2bdb619	Rename all `ParseSess` variables/fields/lifetimes as `psess`. Existing names for values of this type are `sess`, `parse_sess`, `parse_session`, and `ps`. `sess` is particularly annoying because that's also used for `Session` values, which are often co-located, and it can be difficult to know which type a value named `sess` refers to. (That annoyance is the main motivation for this change.) `psess` is nice and short, which is good for a name used this much. The commit also renames some `parse_sess_created` values as `psess_created`.	2024-03-05 08:11:45 +11:00

1 2

97 commits