Concepts to understand before we start reading from a file #
1. The buffer #
A chunk of memory we set aside so that data arriving from somewhere (a file, a network socket, etc) has a home. In the code we’re about to write our buffers will always have a fixed-size.
var buffer: [4096]u8 = undefined;
Notice a few things in the buffer above:
u8is one byte, and since files are just sequences of bytes, a byte array is what we’re going to read them into.[4096]u8means the size is 4096 bytes. This is part of the type and can’t change. No allocation happening here so this buffer lives on the stack.undefinedtells the compiler to reserve the space but not initialize it. This is legal as long as we don’t read the memory before we write to it.
2. Arrays vs. slices #
- An array is fixed in size and known at compile time, it owns it’s storage.
- A slice is a view into an array (a pointer and a length). The pointer points into some area in memory and the length is how much of the memory to view.
Reads rarely fill an entire buffer, the last chunk of a file almost never lines up exactly, so every read tells you how many bytes are valid, and you slice the buffer down to just that part, for example:
const chunk = buffer[0..bytes_read];
After a read we’ll never use the whole buffer, we’ll use a slice of what was actually read. Because slices are a view of underlying memory, when a reader hands you a line as a slice, that slice points into the reader’s buffer, and the next read will likely overwrite it so we’ll need to make sure we use or copy before reading again.
3. The Io value
#
The world outside of your running code is accessible via a value of type std.Io (files, networks, etc). We’re going to build one at the top of main and pass it to every I/O call we make:
const allocator = std.heap.page_allocator;
var threaded: std.Io.Threaded = .init(allocator, .{});
defer threaded.deinit();
const io = threaded.io();
The Io value will help in a few ways:
- Any function that takes an
ioparameter can do I/O, and any function that doesn’t, well, doesn’t. This helps keep function signatures honest. By looking at a function signature alone you can tell what it can and can’t do. - Async event loops, tests, etc, all satisfy the same
std.Io.Threadedinterface, so our file-reading code will work with all of them unchanged.
4. Allocators and stack vs. heap #
The fixed array we talked about above [4096]u8 is very fast (much faster than heap) but it can’t grow and the size must be chosen before the program runs (compile time). When you don’t know how much memory you’ll need, you need the heap, and in Zig all heap memory comes from an allocator that you pass around. Much like our io above any function that allocates passes an allocator in the function signature so you always know where memory is being created. We can always lie, but don’t, it’s bad code.
Let’s write code #
Approach 1: reading a file in fixed-size chunks #
In this approach we ask the operating system for N bytes over and over until it there are no bytes remaining to read. The buffer is reused for every chunk we read, so a large file can be read using only a 4KB buffer on the stack.
const std = @import("std");
pub fn main() !void {
const allocator = std.heap.page_allocator;
var threaded: std.Io.Threaded = .init(allocator, .{});
defer threaded.deinit();
const io = threaded.io();
var buffer: [4096]u8 = undefined;
const file = try std.Io.Dir.cwd().openFile(io, "file.txt", .{});
defer file.close(io);
while(true) {
const bytes_read = file.readStreaming(io, &.{&buffer}) catch |err| switch (err) {
error.EndOfStream => break,
else => return err,
};
const chunk = buffer[0..bytes_read];
std.debug.print("{s}", .{chunk});
}
}
Approach 2: entire file with an allocator #
This approach asks the operating system for exactly as much memory as the file needs to load at runtime. We’ll add a safety cap so we don’t open files that are too large. In practice I prefer approach #1 but this is good knowledge to have.
const std = @import("std");
pub fn main() !void {
const heap_allocator = std.heap.page_allocator;
var threaded: std.Io.Threaded = .init(heap_allocator, .{});
defer threaded.deinit();
const io = threaded.io();
var arena = std.heap.ArenaAllocator(heap_allocator);
defer arena.deinit();
const arena_allocator = arena.allocator();
const file_contents = try std.Io.Dir.cwd().readFileAlloc(
io,
"file.txt",
arena_allocator,
.limited(1024 * 1024),
);
std.debug.print("{s}\n", .{file_contents});
}
Approach 3: read line by line #
In our final approach we’ll process one line at a time. The gotcha here is that the buffer must be at least as large as the longest line we expect to read. Each line handed to us by the reader will contain all the bytes it finds leading up to a newline character, but it can only search within the buffer it has, so a line longer than the buffer (no newline character found and not end of stream) will product a stream too long error.
const std = @import("std");
pub fn main() !void {
const allocator = std.heap.page_allocator;
var threaded: std.Io.Threaded = .init(allocator, .{});
defer threaded.deinit();
const io = threaded.io();
const file = try std.Io.Dir.cwd().openFile(io, "file.txt", .{});
defer file.close(io);
var file_buffer: [1024 * 8]u8 = undefined;
var file_reader = file.reader(io, &file_buffer);
const reader: *std.Io.Reader = &file_reader.interface;
while (try reader.takeDelimiter('\n')) |line| {
std.debug.print("{s}\n", .{line});
}
}