And finally we can implement Pattern machinery for OsStr - #161610
Conversation
The std::sys::os_str::{Buf, Slice} types are only used within the std
crate and not actually exported. Whole `sys` module is private. They
don't need to be public. This might result in a better generated code,
but more importantly it avoids some compile errors down the line.
Firstly, combine functions and results lists into a single list with 'function => result' pairs. This makes it easier to match function with its result. Secondly, eliminate InRange step so that it's easier to notice series of matches or rejects. @pacak: I added a variant to test_stress_indices that matches stuff
Right now things are undertested and underspecified. Some of the library code would get in a loop if searcher starts returning empty rejects. And there's no tests for backwards multi byte char matchers. Pull request I'm reviving had a problem implementing that, so making sure it's tested before the actual code lands. Right now it is possible to break both tests (and user code) without breaking anything else in the test suite I think.
Surprisingly enough there's no rfind tests for multibyte needles, at least it is possible to break this test without breaking anything other test.
Add a Haystack trait describing something that can be searched in and make core::str::Pattern (and related types) generic on that trait. This will allow Pattern to be used for types other than str (most notably OsStr). This somewhat follows the Pattern API 2.0 design. While that design is apparently abandoned (?), it is somewhat helpful when going for patterns on OsStr, so I’m going with it unless someone tells me otherwise. ;) For now leave Pattern, Haystack et al in core::str::pattern. Since they are no longer str-specific, I’ll move them to core::pattern in future commit. This one leaves them in place to make the diff smaller. @pacak: I moved some (or all new) of the `P: Pattern<&'a str> constraints into where clause to keep things narrower: ``` pub fn foo<'a, P: Pattern<&'a str>>(&'a self, pat: P, ...) ... ``` to ``` pub fn replacen<'a, P>(&'a self, pat: P, ...) ... where P: Pattern<&'a str>, ``` Original code had indices in Haystack abstracted as an associated type Cursor. Replaced with usize - Cursor adds noise with not much value. Changed wording in 2-3 places - for example Searcher is generic over a few types so it makes more sense to talk about split points in general with utf8 split points as an example for `&str`.
Pattern is no longer str-specific, so move it from core::str::pattern module to a new core::pattern module. This introduces no changes in behaviour or implementation. Just moves stuff around and adjusts documentation.
Introduce core::pattern::Split and core::pattern::SplitN internal types which can be used to implement iterators splitting haystack into parts. Convert str’s Split-family of iterators to use them. In the future, more haystacks will use those internal types. Co-authored-by: Peter Jaszkowiak <p.jaszkow@gmail.com> @pacak: Fixed some typos, added a few `#[inline]`. Since there's no `H::Cursor` - I had to add `ctx: PhantomData<H>`.
This reverts commit 85cf233ced0d0fe02734c8a83b6d79ccc5432d06. Gone for now, I'll reimplement it later in str_bytes.rs, will confirm with the benchmarks included that the optimization still applies
Introduce core::pattern::EmptyNeedleSearcher internal type which implements logic for matching an empty pattern against a haystack. Convert core::str::pattern::StrSearcher to use it. In future more implementations will take advantage of it. Also adapt and rework TwoWayStrategy into an internal SearchResult trait which abstracts differences between Searcher’s next, next_match and next_rejects methods. It makes it simpler to write a single generic method implementing optimised versions of all those calls. @pacak: - Fixed a few typos. - There's no H::Cursor parameter so code gets a bit simplified. - Added a test to assert how TwoWaySearcher runs with EmptyNeedleSearcher
@pacak: - made more things const fn - there was a (copy-paste?) error in try_finish_byte_sequence so I added a test that checks try_next_code_point(_reverse) with some values, including invalid ones. - reworded a few comments (passive voice, etc) Also different comments: since former is public and later is private due to historical reasons. > This is different than [`next_code_point`] in that it doesn't assume > This is different than `next_code_point_reverse` in that it doesn't assume
Introduce a new core::str_bytes module with types and functions which handle string-like bytes slices. String-like means that they code treats UTF-8 byte sequences as characters within such slices but doesn't assume that the slices are well-formed. A `str` is trivially a bytes sequence that the module can handle but so is OsStr (which is WTF-8 on Windows and unstructured bytes on Unix). Move bunch of code (most notably implementation of the two-way string-matching algorithm) from core::str to core::str_bytes. Note that this likely introduces regression in some of the str function performance (since the new code cannot assume well-formed UTF-8). This is going to be rectified by following commit which will make it again possible for the code to assume bytes format. This is not done in this commit to keep it smaller. @pacak: - Added a few comments - tried to hide internal types from the diagnostic And then there's two different bugs where it would report matched areas as rejected. This broke str::trim_end_matches and who knows what else. Caught it thanks to tests in the previous commit. And one underflow bug on invalid input.
It works right now, but original implementation of the next commit breaks them with none of existing tests catching this regression.
Since core::str_bytes module cannot assume byte slices it deals with are well-formed UTF-8 (or even WTF-8), the code must be defensive and accept invalid sequences. This eliminates optimisations which would be otherwise possible. Introduce a `Flavour` trait which tags `Bytes` type with information about the byte sequence. For example, if a `Bytes` object is created from `&str` it’s tagged with `Utf8` flavour which gives the code freedom to assume data is well-formed UTF-8. This brings back all the optimisations removed in previous commit. @pacak: - removed IS_WTF8 associated constant - unused - fixed a bug related to multibyte reverse matching: `next_code_point_reverse` reads the input via Iterator::next_back, passing `bytes.iter().rev()` reverses it a second time. Not good.
I reverted `ByteNeedle` change earlier, time to add the same functionality back. `ByteSearcherState` is mostly copied from `CharSearcherState`, does a single ascii byte search. It is possible to do the dispatch inside of a CharSearcherState, but that makes it a bit slower.
pattern::find_str 4775.12ns/iter -> 2561.57ns/iter
pattern::rfind_str 5621.05ns/iter -> 2492.68ns/iter
Implement Haystack for &OsStr and Pattern<&OsStr> for &str, char and Predicate. Furthermore, add prefix/suffix matching/stripping and splitting methods to OsStr type to make use of those patterns. Using OsStr as a pattern is *not* implemented. With rust-lang#118485 - I added find and rfind
To work around orphan rules, introduce a wrapper type for predicate
functions to be used as pattern. Specefically, if we want to add
predicat pattern implementation for OsStr type, doing it with a naked
`FnMut` results in compile-time errors:
error[E0210]: type parameter `F` must be covered by another type when it
appears before the first local type (`OsStr`)
impl<'hs, F: FnMut(char) -> bool> core::pattern::Pattern<&'hs OsStr> for F {
^ type parameter `F` must be covered by another type
when it appears before the first local type (`OsStr`)
Due to technical limitations adding support for predicate as patterns on OsStr slices must be done via core::pattern::Predicate wrapper type. This isn’t ideal but for the time being it’s the best option I've came up with. The core of the issue (as I understand it) is that FnMut is a foreign type in std crate where OsStr is defined. Using predicate as a pattern on OsStr is the final piece which now allows parsing command line arguments.
Sadly MultiCharEq needs to be public if we want to keep it in the same place as other related code. Can probably make it sealed. Going via `AsRef<[char]>` comes at about 20% performance penalty
Mirror the single-byte str pattern benchmarks added in the core pattern
benches (find/rfind of a one-byte &str needle) but run them against an
&OsStr haystack, so the byte fast path is exercised through the
str_bytes searcher used for OsStr.
Checking if exposing `MultiCharEq` trait makes sense or we can get away
with just AsRef<[char]>.
```
impl<'hs, F: Flavour, C: AsRef<[char]>> MultiCharEqSearcher<'hs, F, C> {
#[inline]
fn find_match_fwd(&mut self) -> Option<(usize, usize)> {
let mut start = self.start;
while start < self.end {
let (idx, chr, len) = self.haystack.find_code_point_fwd(start..self.end)?;
if self.char_eq.as_ref().contains(&chr) {
return Some((idx, len));
}
start = idx + len;
}
None
}
}
```
pattern::split_char_dense_os 56657.91ns/iter +/- 574.52
pattern::split_char_dense_os_char 56518.98ns/iter +/- 2369.45
pattern::split_char_dense_os_char_arr 75750.77ns/iter +/- 2756.48
pattern::split_char_dense_os_char_arr_ref 74265.01ns/iter +/- 1769.23
pattern::split_char_dense_os_char_slice 83431.05ns/iter +/- 2287.19
implementation with a custom trait
pattern::split_char_dense_os 56535.29ns/iter +/- 6107.48
pattern::split_char_dense_os_char 56390.79ns/iter +/- 789.20
pattern::split_char_dense_os_char_arr 61147.71ns/iter +/- 1891.70
pattern::split_char_dense_os_char_arr_ref 61161.11ns/iter +/- 1088.32
pattern::split_char_dense_os_char_slice 61456.94ns/iter +/- 1537.91
|
cc @Amanieu, @folkertdev, @sayantn Any special-casing of Miri in the standard library requires review. cc @rust-lang/miri |
|
rustbot has assigned @Mark-Simulacrum. Use Why was this reviewer chosen?The reviewer was selected based on:
|
|
I'm aware, will be rebasing/cleaning up as prerequisite get merged. |
|
☔ The latest upstream changes (presumably #161638) made this pull request unmergeable. Please resolve the merge conflicts by rebasing. |
And finally we can add pattern match machinery to
OsStrtype and extend benchmarks to use them.There are a few open questions:
&strgives a way to access patterns with closures (s.strip_prefix(|c| c == 'x')). This works because Pattern, &str and FnMut all living in core. ForOsStryou have to usepredicatefunction (os.strip_prefix(predicate(|c| c == 'x'))) to get around the fact thatOsStrlives instd.AsRef<[char]>- this costs some performance, I have benchmark results somewhere.Contains commits from #161608 #161606 #161604 #161596 #161592 #161589 #160971, for review purposes please look at commits "add pattern matching to OsStr" onwards. I'll will be rebasing stuff as related pull requests get meregd.
This commit is part of the #160971 cinematic universe.