AlgoMaster Logo

Reading & Writing Files

Medium Priority22 min readUpdated June 6, 2026

The System.IO.File class ships with a small set of one-call helpers that handle the most common file work: pull the whole file into a string, write a string back out, append a line to a log, dump bytes to disk. These methods open the file, do the read or write, and close the handle for you, which is why they read so cleanly. This lesson covers what each helper does, when the convenience comes at a real cost (huge files, partial reads, binary records), and how the async variants change the thread story without changing the API shape. Streams, which give you finer-grained control over how bytes move, are the subject of the _Streams_ lesson.

Reading Whole Files

The simplest read in C# is one line. File.ReadAllText opens the file, reads every byte, decodes them as text, and hands you back a single string. It also closes the handle before returning, so you don't have to remember any cleanup.

The first call seeds the file with some sample CSV data so the example is self-contained. The second call reads it back. There's no using block, no Open, no Close. The helper does it for you. That convenience is why these APIs exist: most file work is "give me everything in this file" or "write this string to that file," and most of the time the file is small enough that the trade-offs don't matter.

When the file's contents are line-oriented (a log, a CSV, or a list of product IDs), File.ReadAllLines is usually the better choice. It returns a string[], one entry per line, with the line terminators stripped.

File.ReadAllLines handles the line-splitting using whatever line endings the file uses (\n, \r\n, or \r), and the trailing empty line, if any, is dropped. The returned array is a regular string[], so you can index, slice, sort, or LINQ over it like any other array.

For binary content (an image, a PDF, a serialized blob), the equivalent is File.ReadAllBytes. It returns a byte[] containing the file's raw contents.

The byte values shown happen to be the PNG file signature, the first eight bytes that identify a file as a PNG. Reading those bytes lets you check the format without parsing the whole image. ReadAllBytes is the standard pick for any binary content small enough to fit in memory.

The three "read everything" helpers share one important property: they materialize the entire file in memory before you see any data. For a 4 KB CSV that's a non-issue. For a 4 GB log file, it's a recipe for an OutOfMemoryException.

ReadAllText, ReadAllLines, and ReadAllBytes allocate enough memory to hold the entire file. A 1 GB text file becomes a 1 GB+ string in memory (or 2 GB+ if the file is ASCII and the string is UTF-16, since .NET strings are two bytes per character). For anything larger than a few hundred MB, switch to streaming.

Lazy Reads with File.ReadLines

File.ReadAllLines reads the entire file into a string[] before returning. File.ReadLines (no "All") returns an IEnumerable<string> instead. The signatures look almost identical, but the behavior is fundamentally different.

The file contains a million lines. File.ReadAllLines would allocate a string[] of length 1,000,000 and a million string instances on the heap, just to find the first line ending in "999" and throw the rest away. File.ReadLines reads lines one at a time as the enumerator pulls them, so once FirstOrDefault matches the 999th line, enumeration stops, the file handle closes, and the rest of the file is never read.

The difference is laziness. ReadAllLines is eager: it does all the work up front and returns a populated array. ReadLines is lazy: it returns an iterator that does the work as you ask for each line. For partial reads, scanning, filtering, or any case where you might stop early, the lazy version uses dramatically less memory.

The table below summarizes when to pick which.

MethodReturnsWhen the file is readMemory costGood for
File.ReadAllLinesstring[]All at once before returningWhole file in memorySmall files, when you need an indexable array
File.ReadLinesIEnumerable<string>One line at a time as enumeratedOne line at a timeLarge files, scanning, filtering, partial reads

There is a catch with ReadLines. The file handle stays open until the enumerator is disposed. If you stash the IEnumerable<string> somewhere and forget to enumerate it (or enumerate it twice), the handle behavior can be unexpected. Each enumeration opens and closes the file, so iterating twice reads the file twice. Most of the time you use ReadLines inside a single foreach or LINQ chain, and it works.

The Skip(1) drops the header line, the Count walks through the remaining lines one by one, parsing the third column. At no point does the whole file sit in memory. If the file had ten million product rows, the same code would still use a constant, small amount of memory.

The flow looks like this.

The diagram shows the pull-based shape. The caller asks for one line, the iterator reads one line from disk and yields it, then waits. If the caller stops asking, no more disk I/O happens and the file gets closed. This model is what makes ReadLines cheap on huge files.

File.ReadAllLines on a 1 GB log file allocates roughly 2 GB of managed memory (UTF-16 strings) plus an array large enough to hold every entry. File.ReadLines on the same file holds one line at a time. For anything past 100,000 lines or a few hundred MB, prefer ReadLines.

Writing Whole Files

The write side mirrors the read side. File.WriteAllText takes a string and writes it to a file, creating the file if it doesn't exist and overwriting it if it does. File.WriteAllLines takes a string collection and writes each entry on its own line. File.WriteAllBytes takes a byte array and writes it raw.

The "overwrite if it exists" behavior matters. Both WriteAllText and WriteAllLines truncate the existing file before writing, so if you run them twice with different content, only the second call's content remains. There is no warning, no error, no prompt. To add to the existing content instead of replacing it, use Append, not Write.

The lines variant adds a line terminator after every entry, including the last one. The terminator is whatever Environment.NewLine is on the current platform: \r\n on Windows, \n on Linux and macOS. So a three-line file written on Windows comes out 6 bytes longer than the same three-line file written on Linux for the same content. Most of the time this doesn't matter, but if you're writing files that another tool consumes byte-for-byte, it's something to be aware of.

The first call wrote three product stock levels. The second call replaced the entire file with one line. The original three lines are gone. The directory the file lives in must already exist; if it doesn't, all three writers throw DirectoryNotFoundException. Use Directory.CreateDirectory(folder) before writing if you're not sure the folder is there.

By default, WriteAllText and WriteAllLines write UTF-8 text without a byte-order mark (BOM) in .NET 6 and later. Earlier versions wrote UTF-8 with a BOM by default, which caused interop headaches when other tools choked on the leading bytes. If you're maintaining cross-version code, you can pass an explicit encoding to be sure of the behavior.

Appending Instead of Overwriting

For workloads that grow over time (log files, audit trails, order histories), each call should add to the existing contents rather than replace them. File.AppendAllText and File.AppendAllLines are the append-friendly counterparts. They open the file in append mode if it exists, create it if it doesn't, and add the new content at the end.

Each call adds its content to the end of the file. The first call also creates the file because it doesn't exist yet. AppendAllText does not add a newline for you; the calling code includes \n in each string. AppendAllLines does add a line terminator after each entry, the same way WriteAllLines does.

Each entry lives on its own line. If you call AppendAllLines again, the new entries land after the existing ones, so the file grows over time.

The append helpers are convenient but they're not free for high-frequency writes. Each call opens the file, seeks to the end, writes, and closes. For three log entries that's fine. For a million log entries written one at a time, the open-write-close overhead dominates and a StreamWriter that holds the file open across writes is dramatically faster.

AppendAllText and AppendAllLines open and close the file on every call. For high-frequency writes (logging at thousands of messages per second), that overhead becomes the bottleneck. Use a StreamWriter opened once, written many times, and closed when done.

The decision flow for choosing a write helper looks like this.

Two questions select the helper: text or bytes, and replace or append. Bytes always go through WriteAllBytes and there's no append helper for raw bytes (use a FileStream opened with FileMode.Append if you need it).

Encoding: How Text Becomes Bytes

A file on disk is bytes. A C# string in memory is a sequence of UTF-16 code units. Going between the two requires an encoding, a rule for how characters map to bytes. The text-based helpers use UTF-8 by default, which is the appropriate choice in most cases, but you can override it.

The two files contain the same text, but the BOM-prefixed version is three bytes longer. Those three bytes (EF BB BF) are the UTF-8 byte-order mark, a signature some tools use to detect the encoding. .NET 6 dropped the BOM from the default UTF-8 encoding because most consumers don't want it and it caused bugs (PHP scripts breaking, JSON parsers choking, shell scripts misinterpreting the leading bytes).

The Euro sign in the greeting takes three bytes in UTF-8 (E2 82 AC). Other multi-byte characters (Chinese, Japanese, emoji) take three or four bytes each. ASCII characters (a-z, A-Z, 0-9, basic punctuation) take exactly one byte. UTF-8 is variable-width, which is why "the byte length of a string" is not the same as "the character count of a string."

If you write the same greeting in ASCII, you get an exception: ASCII can't represent the Euro sign at all. The encoder either throws or silently substitutes a replacement character, depending on the encoder's fallback configuration.

ASCII is fixed-width: every character is exactly one byte. The 32-character string becomes 32 bytes on disk. That predictability is why some legacy file formats and protocols still mandate ASCII, but for general-purpose text in a modern app, UTF-8 is the standard default.

The same encoding parameter exists on ReadAllText, ReadAllLines, ReadLines, and the Async variants. If you wrote a file with a non-default encoding, you must read it back with the same encoding, otherwise the decoder produces garbage (or throws, depending on settings).

The accented é was written as a single byte using the Latin-1 (ISO-8859-1) encoding. Reading it back with the same encoding restores the original. Reading it as UTF-8 fails to decode that byte and produces a replacement character. The fix is to use UTF-8 everywhere unless an external system forces something else.

EncodingBytes per charRangeDefault in .NET 6+
Encoding.UTF81-4 (variable)All UnicodeYes (no BOM)
Encoding.ASCII1 (fixed)First 128 code pointsNo
Encoding.Unicode (UTF-16 LE)2-4All UnicodeNo
Encoding.UTF324 (fixed)All UnicodeNo
Encoding.GetEncoding("ISO-8859-1")1 (fixed)Latin-1 (256 chars)No

The rule of thumb: write UTF-8 by default, only change it when something downstream requires it, and always read with the same encoding you wrote with.

Async Variants

Every read and write helper has an Async counterpart: ReadAllTextAsync, WriteAllTextAsync, ReadAllLinesAsync, WriteAllLinesAsync, ReadAllBytesAsync, WriteAllBytesAsync, AppendAllTextAsync, AppendAllLinesAsync. The shape of the call is the same, just with await and a Task-returning signature.

The two-line read-and-write pattern from earlier in the lesson is the same here, with Async suffixes and await keywords. The _Async Programming Basics_ lesson covered the why of async: I/O-bound operations release the calling thread during the wait instead of blocking it, which matters in web servers, UI apps, and anywhere thread reuse is important.

Two things to know about the async file helpers.

First, file I/O is a mix of true asynchronous I/O at the kernel level and synchronous-with-fake-async at the .NET level. Whether a given ReadAllTextAsync call genuinely releases the thread for the whole read depends on the platform, the file size, and how the underlying FileStream was configured. For small files on modern .NET, you generally do get real async behavior. For tiny files, the overhead of the async state machine can make the call slower in wall-clock time than the sync version, even as it frees the thread.

Second, the async variants accept a CancellationToken, which the sync variants do not. That means a long-running file write can be cancelled cooperatively if the caller passes a token and trips it.

Output (one of):

or

The output depends on whether the write finishes before the 50 ms timer fires. If the file is on a fast SSD, 100 MB might write before the cancellation triggers. On slower storage, the cancellation succeeds and the partially written file is left on disk. Cancellation tokens are covered in detail in the _CancellationToken_ lesson.

The decision rule for sync vs async file helpers is the same as for any I/O. In a web server, ASP.NET handler, UI app, or anything else where thread reuse matters, use the async variants. In a one-shot console script that does its work and exits, the sync variants are fine and a bit simpler to read.

The async variants compose with Task.WhenAll for parallel file work, which is one of the cases where async changes the wall-clock time, not just the thread cost.

Task.WhenAll starts all three reads concurrently and resolves when all three finish. If each read takes 20 ms, the total wall-clock time is about 20 ms (the longest individual read), not 60 ms. Task.WhenAll is covered in detail in the _Task.WhenAll & Task.WhenAny_ lesson.

Async file helpers have a small per-call overhead (state machine allocation, continuation bookkeeping). For a single 1 KB file, the sync version is faster. For a 10 MB file, the costs are negligible relative to the I/O. For a web server doing this on every request, the async version's thread-pool savings are what matter, not the per-call timing.

When These Helpers Are the Wrong Choice

The convenience APIs cover the common case well. They also have a few situations where they're a bad fit, and the answer is a Stream or a BinaryReader/BinaryWriter. Knowing what those situations look like helps you pick the appropriate API from the start.

Huge files. "Huge" depends on your environment, but as a working rule, anything over a few hundred MB is a candidate for streaming instead of ReadAllText/ReadAllBytes. The convenience APIs allocate enough memory to hold the entire file. A 4 GB log file becomes a 4 GB+ string in memory. On a 32-bit process or a memory-constrained container, this is an instant crash. On a 64-bit process with plenty of RAM, it still wastes garbage collector time and can trigger large-object-heap fragmentation. File.ReadLines handles huge text files lazily, but for huge binary files, you want FileStream and a buffered read loop.

Partial reads. If you only need bytes 1024 through 2048 of a file, ReadAllBytes is wasteful: it reads the whole file just so you can grab a slice. A FileStream lets you Seek to a specific position and read just the bytes you need. This shows up in formats like ZIP files, image headers, or any binary structure with an index at the front.

Binary records with a known structure. If a file contains, say, 1000 fixed-size product records of 64 bytes each, ReadAllBytes gives you the bytes but not the structure. BinaryReader (lesson 04) lets you read primitives directly: ReadInt32, ReadDouble, ReadString. The code stays close to the file format, instead of you doing manual byte arithmetic on a byte[].

Streaming output. When you're building a response on the fly, like generating a CSV from a database query without ever materializing the full result in memory, you want to write rows as they come, not collect them all in a string[] and call WriteAllLines. StreamWriter lets you write lines incrementally while holding the file open across writes.

High-frequency writes. A logger writing thousands of lines per second through AppendAllText will be bottlenecked on the open-write-close cycle. A StreamWriter opened once at startup, written through over the app's lifetime, and flushed periodically, is the appropriate shape.

Concurrent access from multiple processes. The convenience APIs use FileShare.Read by default, which means another process can read while you're writing but the semantics get messy fast. For files that need careful sharing semantics (databases, log rollover, lock files), explicit FileStream construction with the appropriate FileShare and FileAccess arguments is the way to go.

The pattern across all six is the same: any time you need fine-grained control over how the bytes move (which range, which order, when the file opens and closes, who else can touch it), the convenience APIs hide too much. They were designed for "give me everything" or "save this thing," not for orchestrating I/O.

SituationWrong toolRight tool
4 GB log file scanReadAllText (out of memory)File.ReadLines or FileStream
4 GB binary blobReadAllBytesFileStream + buffered reads
Reading bytes 1024-2048ReadAllBytes then sliceFileStream + Seek
Parsing 1000 fixed-size product recordsReadAllBytes + manual byte mathBinaryReader
Generating CSV row by rowWriteAllLines (collects everything first)StreamWriter line by line
10K log lines per secondAppendAllText per lineStreamWriter held open
Multiple processes writingDefault WriteAllTextFileStream with explicit FileShare

For everything else, File.ReadAllText, File.WriteAllText, File.ReadLines, File.AppendAllText, and their async cousins are the appropriate APIs and there's no reason to go lower-level.

Quiz

Reading & Writing Files Quiz

10 quizzes