XML serialization is the process of turning a C# object into XML text and reading it back into an object. JSON has taken over most modern APIs, but XML is still everywhere: SOAP services, build files like .csproj, configuration formats like web.config, document-centric data exchange, and any system that wants schema validation through XSD. This lesson covers the two main APIs that ship with .NET, XmlSerializer for object mapping and the XmlReader/XmlWriter/XDocument family for direct XML manipulation, plus the trade-offs between them and the pitfalls that arise when the wrong tool is picked.
XML is not the trendy format in 2026, but a survey of the .NET ecosystem makes it clear it is not going anywhere. Every .csproj, .sln, and .config file you have ever opened is XML. NuGet's package metadata (.nuspec) is XML. The classic ASP.NET configuration system (web.config) is XML, and even modern .NET still reads app.config for some scenarios. Build engines like MSBuild are entirely XML-driven. Documentation comments compile down to XML files that IDEs read for IntelliSense.
Outside the .NET world, XML is the language of integration with older systems. SOAP web services, which are still common in finance, healthcare, government, and enterprise software, send XML envelopes over HTTP. EDI gateways, banking message formats like ISO 20022, and many B2B exchanges all speak XML. If you write code in any of those domains, you will read and write XML whether you want to or not.
XML also has features JSON does not. A schema language (XSD) lets you validate documents against a contract before processing them. Namespaces let two XML vocabularies live side by side in the same document without colliding. Attributes let you attach metadata to elements without inflating the structure. XPath and XSLT let you query and transform documents in ways that JSON tooling has been slowly catching up to. None of these features are "better than JSON" in general, but each one fits some real problem.
The good news is that .NET has had first-class XML support since version 1.0. The APIs are stable, well-documented, and battle-tested. The bad news is that there are several of them, and they overlap in confusing ways. The rest of this lesson sorts them out.
The diagram captures the round trip. An object goes through a serializer to become XML text, the text is written to a file or stream, and the same path runs in reverse to load the object back. Both the object-mapping API (XmlSerializer) and the document API (XDocument, XmlWriter) sit in the middle box. The choice between them is the main design decision when you work with XML in C#.
XmlSerializer, in System.Xml.Serialization, is the high-level API. Pass it a C# object and it produces XML; pass it XML and a target type, and it builds the object back. It uses reflection on the type the first time it sees it, generates a serialization assembly in memory, and then runs that assembly for every subsequent call. The result is fast steady-state performance, but a noticeable cold-start cost the first time a new type is serialized.
The basic pattern is short.
A few defaults to know. The root element is named after the class (<Product>). Each public read-write property becomes a child element with the same name. Two XML schema namespaces (xsi and xsd) are added to the root by default, even when nothing in the document uses them. The XML declaration says utf-16 because we wrote to a StringWriter, which is UTF-16 in memory. Writing to a file with the appropriate encoding produces a more useful declaration.
XmlSerializer has two hard requirements on the type it serializes. First, the type must have a public parameterless constructor. The deserializer creates the object with new T() and then sets each property from the XML. If your class only has a constructor that takes arguments, deserialization throws InvalidOperationException with a message about the missing constructor. Second, only public read-write properties and fields are serialized. Private members, read-only properties without setters, and indexers are skipped silently.
Reading XML back into an object is the mirror of writing.
Deserialize returns object, so a cast to the expected type is required. The ! after the cast is the null-forgiving operator; Deserialize can in theory return null for an empty stream, and the compiler warns about the cast unless you tell it the value is not null.
Constructing a new XmlSerializer(typeof(T)) for a type the runtime has never seen does heavy work: it generates a serialization assembly via reflection. The first call for a given type can take tens of milliseconds. Subsequent calls reuse the cached assembly and are fast. In production, cache the XmlSerializer instance per type (a static readonly field works) instead of creating one on every call.
The default mapping is fine when the C# class shape and the XML shape match exactly. Real systems rarely cooperate that nicely. You usually need to rename elements, move some values to attributes, ignore certain properties, or wrap a list in a parent element. System.Xml.Serialization ships a set of attributes that control all of this. The most common ones are listed below.
| Attribute | What it does |
|---|---|
[XmlRoot("name")] | Renames the root element of the document. |
[XmlElement("name")] | Renames a property's element, or controls collection element names. |
[XmlAttribute("name")] | Serializes the property as an XML attribute on the parent element instead of a child element. |
[XmlIgnore] | Skips a property entirely; it is not written or read. |
[XmlArray("name")] | Renames the wrapper element around a collection. |
[XmlArrayItem("name")] | Renames each item element inside a collection. |
[XmlText] | Marks a property to receive the inner text of the element rather than a child element. |
[XmlEnum("value")] | Maps an enum member to a different XML name. |
[XmlNamespaceDeclarations] | Lets the object expose its own namespace declarations. |
[XmlInclude(typeof(Derived))] | Tells the serializer about a derived type so polymorphic references can be serialized. |
A concrete example shows how these compose. The order class below uses attributes for identity fields and elements for content.
The structure: id and currency are attributes on <Order>, not child elements. The customer name is a child element renamed from CustomerName to <Customer>. The list of lines is wrapped in <Items> thanks to [XmlArray], and each line is named <Item> thanks to [XmlArrayItem]. The LoadedAt property is marked [XmlIgnore] and does not appear at all.
The general rule for choosing between an attribute and an element is taste plus convention. Identifiers, types, and metadata that classify an element typically go on attributes. Content data, especially anything that might grow or contain nested structure, goes in child elements. There is no compiler enforcement; pick a convention and stick to it across the document.
By default, XmlSerializer adds two namespace declarations to the root element: xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" and xmlns:xsd="http://www.w3.org/2001/XMLSchema". These are useful when your XML actually uses XSD-typed values or xsi:nil, but most of the time the document does not use them and they just clutter the output. They also waste bytes on every payload, which adds up on a busy SOAP endpoint.
The fix is to pass an XmlSerializerNamespaces instance to Serialize. An empty one suppresses the defaults; a populated one declares your own namespaces.
When you do need a real namespace, declare it on both the type (with [XmlRoot] or [XmlType]) and the serializer call.
The namespace prefix o is bound to the URL once at the root, and every element from that namespace is prefixed accordingly. Picking a sensible prefix matters for readability, especially in documents that mix two or three namespaces.
In real code, you almost never want to round-trip through a StringWriter. You want to write directly to a file or a network stream, and you want to control the encoding. XmlWriterSettings is the lever for both.
Some decisions to call out. UTF8Encoding(false) writes UTF-8 without a byte-order mark (BOM). The BOM is the three-byte sequence EF BB BF at the start of a file that some tools add to mark the encoding. It is harmless for many parsers, but other tools (especially older Java and Python XML libraries) treat it as garbage at the start of the document and fail. Omitting the BOM is the safer default for interop. Indent = true produces human-readable XML; turn it off for compact wire formats. OmitXmlDeclaration = false keeps the <?xml version="1.0" encoding="utf-8"?> declaration; some consumers reject documents without it, others reject documents with it, so you have to know which side you are talking to.
The XmlReaderSettings counterpart controls reading: schema validation, DTD handling, ignoring whitespace and comments, and a few security knobs.
DtdProcessing.Prohibit is the safer default for any XML coming from outside your own systems. DTDs can encode entities that expand to gigabytes of text or trigger billions of entity lookups, the classic "billion laughs" attack. Prohibiting DTDs upfront is the simplest defense.
Setting Indent = true adds bytes (newlines and spaces) to every payload. For internal storage and debugging, the readability matters. For high-volume wire formats, leave it off and pretty-print only when humans need to look at the document.
XmlSerializer is convenient but heavy. It assumes you have a class hierarchy that maps to the XML, it builds the entire object in memory, and it pays a one-time reflection cost per type. When you need to stream through a large XML file, or when you just want to write a few elements without modeling them as classes, XmlReader and XmlWriter are the lower-level alternatives.
Both are forward-only. XmlReader walks the document one node at a time, consuming it as it goes. You cannot back up. XmlWriter writes nodes in document order, and once you have closed an element, you cannot go back and add an attribute to it. The trade-off is performance: forward-only parsing uses constant memory regardless of document size, and forward-only writing avoids buffering the whole document in a DOM tree.
A simple XmlWriter example writes the same order document by hand.
The result is byte-for-byte identical to what XmlSerializer would produce for the same data, but the code never built an Order object. For a one-off conversion or a code path that already has the values in local variables, this can be more direct than defining a class with attributes just to satisfy the serializer.
XmlReader walks a document the same way, exposing Read() to advance to the next node, and properties like NodeType, Name, Value, and MoveToAttribute(...) to inspect the current position.
The reader scans through the document once, in order. For a 5 GB XML log file, this approach uses kilobytes of memory regardless of file size. The same job with XmlSerializer or XDocument would try to load the whole document into memory and almost certainly run out.
The cost of all that efficiency is verbosity. You write the parsing logic by hand, including handling nested elements, optional attributes, and edge cases like elements you do not care about. For small documents, the savings are not worth the extra code.
The middle ground between XmlSerializer (object mapping) and XmlReader (raw streaming) is XDocument, the LINQ to XML API in System.Xml.Linq. It loads the whole document into a tree of XElement and XAttribute objects, but the API is built for query and modification rather than serialization. If you need to read a value out of an XML file, change a few elements, or build a small document by hand, XDocument is usually the simplest option.
Building a document looks almost like the structure of the XML itself.
The constructor calls nest exactly like the XML, which makes the structure obvious at a glance. Querying is just as direct.
The casts: (string)item.Attribute("sku")! converts the attribute to a string; the same explicit-conversion pattern works for int, decimal, DateTime, and other built-in types. The cast returns null if the attribute or element is missing, which makes optional values easy to handle with ??. The whole API works like regular C# objects, with LINQ available for filtering, projection, and aggregation.
There is one more older API in the box: XmlDocument, in System.Xml. It implements the W3C DOM specification (the same shape as document.getElementById in JavaScript), and it has been in .NET since version 1.0. New code should use XDocument instead. The DOM API is more verbose, lacks LINQ integration, and uses a less type-safe model. The only reason to use XmlDocument today is interop with older code that already uses it.
| API | Best for | Memory | API style |
|---|---|---|---|
XmlSerializer | Round-tripping classes to XML and back | Whole object in memory | Attribute-driven mapping |
XmlReader / XmlWriter | Streaming through huge documents, performance-critical writes | Constant, regardless of document size | Forward-only, low-level |
XDocument (LINQ to XML) | Querying, modifying, or building small to medium documents | Whole document tree in memory | Functional, LINQ-friendly |
XmlDocument (DOM) | Interop with legacy .NET code that already uses it | Whole document tree in memory | W3C DOM, verbose |
XmlSerializer does not magically know about your derived types. If a property is declared as a base class but the actual instance is a subclass, the serializer needs to be told about the subclass at the point it builds its serialization assembly.
Output (abridged):
The serializer emits an xsi:type attribute on each element to record the actual type. On deserialization, it reads that attribute and constructs the right subclass. Without [XmlInclude], the serializer throws an exception complaining that it does not know how to serialize the derived type.
XmlSerializer is not the only option. DataContractSerializer, in System.Runtime.Serialization, was introduced with WCF (Windows Communication Foundation) in .NET 3.0 and uses an opt-in model: you mark the type with [DataContract] and each member you want serialized with [DataMember]. Members without [DataMember] are ignored, which is the opposite of XmlSerializer's opt-out approach.
The two serializers differ in a few important ways. DataContractSerializer always uses a namespace (defaults to http://schemas.datacontract.org/2004/07/... if you do not set one). It does not support attribute-style XML; everything is elements. It serializes private fields and properties when marked with [DataMember], which XmlSerializer cannot do at all. It is faster and produces smaller payloads in many cases. It also uses an opt-in model, which makes it harder to accidentally serialize sensitive fields.
For new code that controls both ends of the wire and only needs to round-trip C# objects, DataContractSerializer is often a better choice than XmlSerializer. For interop with existing XML schemas where you need full control over element names, attributes, and document shape, use XmlSerializer. Most codebases use one or the other consistently rather than mixing them.
For new internal APIs and modern systems, JSON is almost always the appropriate choice. It is smaller, faster to parse, easier for JavaScript and Python to consume, and most modern HTTP libraries default to it. XML fits specific scenarios.
| Aspect | XML | JSON |
|---|---|---|
| Payload size | Larger (verbose tags, closing elements) | Smaller |
| Parse speed | Slower in most benchmarks | Faster |
| Schema validation | Built-in via XSD, mature tooling | JSON Schema exists but is less standardized |
| Attributes vs elements | Both, with semantic distinction | Object keys only |
| Namespaces | First-class support | Not supported |
| Comments | Supported | Not supported in standard JSON |
| Mixed content (text and elements) | Natural fit | Awkward |
| Tooling in .NET | XmlSerializer, XDocument, XmlReader, XSD generation | System.Text.Json, Newtonsoft.Json |
| Common in modern web APIs | Rare | Default |
| Common in legacy enterprise / SOAP | Required | Rare |
The decision is mostly about who you are talking to. Talking to a SOAP service or a system that already uses XSD for validation, use XML. Building a new HTTP API or saving structured data to disk for your own use, use JSON. Reading a project file or a configuration document, you have no choice; the format dictates the API.
The flowchart captures the practical decision. If any of the three "yes" branches apply, XML is the appropriate choice and one of the .NET XML APIs handles it. Otherwise JSON wins by default for new code.
A short tour of the mistakes that show up most often in production XML code.
Reflection cost on every call. Creating new XmlSerializer(typeof(T)) is expensive the first time for a given type. Every subsequent call for the same type is fast because the serialization assembly is cached, but only if you reuse the same serializer instance. Code that constructs a new XmlSerializer inside a method called on every request pays the assembly-generation cost over and over. The fix is to cache the instance, typically in a static readonly field.
Missing parameterless constructor. If your class only has constructors that take parameters, XmlSerializer.Deserialize throws at runtime. Add a public parameterless constructor, even if you do not want callers to use it. You can sometimes mark the parameterless constructor protected or private and put a [XmlSerialization] attribute, but the simplest fix is to make it public and rely on convention.
XML injection. Building XML by string concatenation is the same kind of mistake as building SQL by string concatenation. If a value contains characters like <, >, or &, those characters can change the structure of the document. Always use the appropriate API (XmlSerializer, XDocument, or XmlWriter) which escapes content correctly. Never write XML like "<Name>" + userInput + "</Name>".
Encoding and BOM mismatches. The XML declaration says one encoding, the actual byte stream uses another, and the parser refuses to load the document. The fix is to be explicit about encoding on both ends. Pass XmlWriterSettings { Encoding = ... } and use a matching StreamReader or stream encoding when reading. UTF-8 without BOM is the most portable choice for cross-platform interop.
Polymorphism without `[XmlInclude]`. A property typed as a base class can hold any subclass at runtime, but XmlSerializer rejects unknown derived types unless they are declared via [XmlInclude] (or passed in the extraTypes argument to the constructor). Without that, you get a runtime exception that is easy to miss in unit tests if your tests only use the base type.
Whitespace surprises. XML preserves whitespace inside elements unless you tell it not to. A document like <Name> Mouse </Name> deserializes as " Mouse ", with the spaces. If you want to trim, do it explicitly in the property setter or after deserialization. XmlReaderSettings.IgnoreWhitespace = true only ignores whitespace between elements, not inside them.
Deserializing untrusted XML. Any time you parse XML from outside your system, set DtdProcessing = DtdProcessing.Prohibit and disable XML resolver behavior. The default settings on modern .NET versions are reasonably safe, but on older versions and on XmlDocument you have to opt into the safe defaults explicitly.
10 quizzes