Skip to content

ACP: Add BufReader::chars() -> Chars #881

Description

@Hexorg

Proposal

Problem statement

I need to read a large technically text file that potentially has new lines, but it's more like a markup of text and reading it line by line is wasteful, breaks apart markup structure, and line presence is not guaranteed. Ideally I'd read it character by character as an iterator - making a streaming shift-reduce parser.

Motivating examples or use cases

Ideally, we can have

let file = File::open("big_file").expect("Unable to open file.");
let mut reader = BufReader::new(file);
let mut parser = Parser::new();
for c in reader.chars() {
    parser.shift_reduce(c)
}

Solution sketch

// std/src/io/mod.rs
// ...
// I looked at str::Chars and I suspect we can't just return it here, because we need to 
/// load the next file chunk once we consume all std::Chars of the contained buffer. 
pub struct Chars<R> {
    inner: R
}
// ...
pub trait BufRead: Read {
        pub fn chars(self) -> Chars<Self> {
               Chars(self)
        }
}
// ...
impl Iterator for Chars<R> where R:BufRead {
    type Item = Result<char>;

    fn next(&mut self) -> Option<Self::Item> {
        // if the buffer is empty, fill it up
        // read buffer as UTF8. If invalid code point return Err() and go to  the next byte
        // otherwise return Ok(char_
       // if reached the end of buffer, refill it. 
        // if EOF return None forever.
    }
}

Alternatives

I've looked around and I didn't find any crates that implement anything like this. The functionality seems to be a bit too simple to turn it into a standalone crate.

Right now BufReader provides buffer() -> &[u8], which I can convert to Utf8Chunks, but none of the mutable buffer functionality is exposed, so I cant fill_buffer as needed by the Utf8Chunk data. Module buffered is private too so I can't make my own BufCharReader. To parse arbitrary UTF8 text, we need at least a 4-byte buffer, so a BufReader seemed a perfect fit for this.

Links and related work

In C++ Visual Studio There's a way to stream utf8 chars, but for GCC this requires a 3rd party library.
Anything remotely-related I found in Rust Internals is about UTF8 prefix inspection, which can somewhat allow this ACP to use a 4-byte buffer, but since BufReader's default buffer size is 8Kb, I think it's fine to use reader.buffer().utf8_chunks() to implement this.

Activity

  1. added
    api-change-proposalA proposal to add or alter unstable APIs in the standard libraries
    on Sep 17, 2026
  2. pitaj commented on Sep 17, 2026

    @pitaj

    Should probably go on the BufRead trait instead, like lines.

    Also please add the API of Chars. Presumable it's an iterator, but what is the element type? etc

  3. Hexorg commented on Sep 17, 2026

    @Hexorg
    Author

    @pitaj I've updated the original text. Thank you.

  4. scottmcm commented on Sep 17, 2026

    @scottmcm
    Member

    It's not obvious to me that getting chars is really something to encourage here. Notably, if you're doing a lexer or parser, as recently discussed on URLO you're probably better off processing the bytes directly rather than lexing to UTF-32 first: Matching the 3 bytes that make up a ∛ is no different from matching the 3 bytes that make up for.

    I wonder about maybe doing something simpler than an iterator to start with. For example, we already have BufReader::peek, so maybe some kind of BufReader::peek_str could be useful? The edge cases there aren't super obvious to me, though -- I guess it's fine to just error if it's invalid UTF-8 (people can handle the UTF-8 conversion themselves if they want partially-UTF-8) but if it's in the middle of a USV should it try to add more to the buffer? Shorten the returned amount to the complete USVs? Have peek_str(N) just do peek(N+3) internally? Dunno.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions