Reflection inspects your code at runtime: it walks types, reads attributes, and builds delegates on the fly. Source generators do the same kind of work, but at compile time, by writing extra C# files into your build before it produces an assembly. This lesson covers what source generators are, the incremental generator pipeline, and a small worked example that emits a ToString() override from a [ToString] attribute. The payoff is the same boilerplate you'd write by hand or fish out with reflection, with zero runtime cost and full compatibility with NativeAOT and trimming.
A reflection-based ToString() helper has a familiar shape. You walk the object's public properties with GetProperties(), read each value, and stitch the names and values into a string. Every call pays for the metadata walk, every property read goes through a PropertyInfo.GetValue call, and the JIT has no way to inline any of it. Under NativeAOT or aggressive trimming, the reflection might not even work, because the trimmer doesn't know which properties survive.
A sketch of the reflection-based version makes the comparison concrete:
The output matches what the source generator produces: the two approaches solve the same problem. The difference is what's happening underneath. GetProperties returns a PropertyInfo[] (allocated, cached, but still a metadata lookup); GetValue(obj) boxes the result of every value-type property into an object; string.Join allocates the joined string; the LINQ Select adds a chain of iterator allocations. None of that shows up in the generator-emitted version, which writes the property values into the interpolated string holes directly.
A source generator flips this around. You declare your intent with an attribute, the generator runs as part of dotnet build, reads the marked types out of the compilation, and emits a new .cs file containing the equivalent hand-written code. By the time your program runs, there is no reflection. There is just a normal ToString() method that the JIT inlines like any other.
The contrast is in where the work happens. Reflection pays at runtime, on every call, in every process, on every machine. The generator pays once during the build that produces your binary. Everyone who runs that binary gets the generated code with no metadata walks and no boxed property values.
A reflection-based formatter that loops over PropertyInfo.GetValue boxes every value-type property on every call. A generator-emitted ToString() writes the values directly into a StringBuilder, no boxing. On a hot logging path that formats thousands of objects per second, that difference shows up as both lower CPU and lower allocations.
A source generator is a small program that the C# compiler hosts during a build. It talks to the compiler through Roslyn, the .NET compiler platform. Three pieces matter. The compilation is the set of all source files, references, and options that make up your build. A syntax tree is the parsed form of one source file, the shape of the code with no meaning attached. The semantic model is the layer that gives those shapes meaning: which Customer does Customer customer refer to, what type does this method return, which attributes are applied to this class. A generator uses these three pieces to find candidate types, decide what to emit, and produce new source files.
You don't need to master Roslyn to write a useful generator. The incremental generator API exposes high-level helpers that hide most of it. The example in this lesson uses one of those helpers and never touches a raw syntax tree. The high-level picture is enough: the compiler parses your code, gives each declaration a typed identity (a Symbol), and offers your generator a chance to ask "any class marked with this attribute? Emit something based on what you see."
When you write Customer customer = new Customer(); in your editor, three things happen at compile time. The lexer breaks the line into tokens. The parser builds a syntax tree where Customer and customer are just names in a tree. The binder runs the semantic model and decides that the first Customer refers to a particular INamedTypeSymbol in your project (and not, say, a Customer in some other namespace). Generators get to read those symbols, which means they see your code the way the compiler does, not the way a text-processing tool would.
There are two generator APIs. The older ISourceGenerator runs end-to-end on every build, which is fine for tiny generators but starts to bite on large solutions because the work isn't cached. The modern IIncrementalGenerator, introduced with .NET 6 tooling, expresses the generator as a pipeline of cached steps. Each step's output is only recomputed when its input changes, so an IDE that runs the generator on every keystroke pays only for the parts of the pipeline whose inputs actually changed.
The pipeline has three moving parts that appear in almost every incremental generator:
| API | Role |
|---|---|
RegisterPostInitializationOutput | Emit fixed scaffolding that doesn't depend on user code (an attribute declaration, a common base type) |
SyntaxProvider.ForAttributeWithMetadataName | Find types in the compilation that carry a given attribute, producing a stream of candidates |
RegisterSourceOutput | For each candidate, generate a new .cs file and add it to the compilation |
The order matters. The generator first emits the marker attribute (so users can apply it), then asks the compilation for every type marked with that attribute, then writes a generated file for each one.
[ToString]The example is a generator that emits a ToString() override on any class marked with [ToString]. The override prints the class name and every public property in the form ClassName { Prop1 = value1, Prop2 = value2 }. Records already provide this, but [ToString] lets you opt regular classes into the same behavior without writing the method by hand.
The full setup is three files: the generator project that hosts the generator and emits the marker attribute, a consumer project that references the generator and uses the attribute, and the generated .cs file the build produces. Take them one at a time.
The generator itself ships as a small .NET Standard library that the consumer references with a special OutputItemType="Analyzer" tag. The project file pulls in Microsoft.CodeAnalysis.CSharp, which is where the incremental generator types live.
A couple of points deserve emphasis. The target framework is netstandard2.0, not net8.0. The compiler runs generators in its own process, and that process expects analyzers compiled against netstandard2.0. Targeting net8.0 here will fail to load. The IsRoslynComponent flag and EnforceExtendedAnalyzerRules exist so the SDK treats this project as a generator and applies the appropriate analyzer rules.
The marker attribute can live anywhere, but the standard approach is for the generator to emit it through RegisterPostInitializationOutput. That way the consumer doesn't need to define [ToString] themselves; it appears as part of the build.
The Initialize method wires up the three pipeline steps. RegisterPostInitializationOutput runs once per compilation and adds the attribute declaration as ToStringAttribute.g.cs. SyntaxProvider.ForAttributeWithMetadataName does the central work: it finds every class declaration in the compilation that carries [ToString], projects it into an INamedTypeSymbol (the semantic-model representation of the type), and filters out any nulls. RegisterSourceOutput is where the new source actually gets emitted; for each candidate symbol, BuildToString produces a string, and AddSource hands it back to the compiler as a new file in the compilation.
Incremental generators rely on the equality of values flowing through the pipeline. If you return a custom record from a pipeline step and that record's Equals doesn't reflect what the downstream code cares about, the cache will think every build has new inputs and rerun the whole pipeline. The standard fix is to use small record types with value equality or to project to the simplest data possible (a string, a tuple of primitives) before RegisterSourceOutput.
A few details inside BuildToString deserve attention. classSymbol.ContainingNamespace.ToDisplayString() produces the fully qualified namespace as a string, exactly the form you'd type at the top of a hand-written file. classSymbol.Name is the unqualified class name. Together they let the generator place the partial declaration into the same namespace and class as the user's original code, which is what makes the merge work.
The loop over classSymbol.GetMembers() walks every member declared on the class. IPropertySymbol is the Roslyn type that represents a property declaration; if the member happens to be a field, a method, or a nested type, it has a different symbol type and falls through the pattern. The accessibility check (DeclaredAccessibility: Accessibility.Public) and the static check (IsStatic: false) are how the generator narrows the property set to the same shape a reflection-based helper would walk. The doubled braces ({{ and }}) in the interpolated string are the way to embed literal braces into a string that's going to become an interpolated string in the generated code; the inner {{prop.Name}} becomes {Id} in the output, which the C# compiler then interprets as a property reference inside its own interpolated string.
This is a deliberately tiny generator. A production-quality version would handle nullable annotations, indexers, properties with bad accessors, classes inside generic outer types, and a few other edge cases. The shape stays the same; the body of BuildToString just gets longer.
The consumer references the generator like an analyzer, not like a normal library. In the consumer's .csproj:
The OutputItemType="Analyzer" tells MSBuild to load the project as an analyzer, which is the same channel source generators use. ReferenceOutputAssembly="false" keeps the generator's compiled DLL out of the consumer's runtime references, because the generator only runs at build time and you don't want its types polluting your runtime API.
With the generator wired in, the consumer code looks like:
The Product class is declared partial so the generator can add the ToString() override in a separate generated file without touching the original source. The [ToString] attribute marks the class for the generator. Everything else looks like ordinary C#.
The generator emits a file the compiler treats as a regular part of your project. The file it produces for Product:
Generated `Product.ToString.g.cs`:
This is hand-written C# in every meaningful sense. The compiler typechecks it, the JIT compiles it, and the call site does a direct virtual dispatch into it. No PropertyInfo, no boxing, no metadata walk. If you renamed Product.Price to Product.Cost, the generated file would change on the next build and the rename would just work.
A generated method can never appear in the same physical file as the user's declaration. The generator emits a separate file (Product.ToString.g.cs in the example), and both files declare the same partial class. The compiler then merges them into a single type at build time. This is the same merging the compiler does for hand-written partial classes; the generator provides one of the parts.
Generated files normally sit in memory during compilation and don't end up on disk. That's a problem when you're debugging a generator or trying to understand what code the build is actually producing. Two MSBuild properties control this:
With EmitCompilerGeneratedFiles set to true, every generated file is written to disk under obj/generated/<Generator.Assembly>/<Generator.FullName>/. You can open those files in any editor, set breakpoints in them (most IDEs handle this automatically once the files exist on disk), and verify that what the generator produced matches what you intended.
Most IDEs also let you navigate into generated code directly. In Visual Studio and Rider, "Go to Definition" on a generator-produced member takes you to the generated file; the file shows up in the Solution Explorer under the project, in a "Dependencies" or "Source Generators" node. JetBrains Rider exposes the same tree under "Compile-time generated files".
A workflow that pays off: turn EmitCompilerGeneratedFiles on while you're authoring the generator, watch the output files appear and update as you iterate, then turn it back off (or leave it on, since the cost is small) once the generator stabilizes. Reading the actual output is the fastest way to catch the kinds of bugs that produce odd-looking but technically-compilable code, like a stray comma in an interpolated string or a missing semicolon at the end of a generated statement.
Leaving EmitCompilerGeneratedFiles on permanently isn't harmful, but it does mean every build writes generated files to disk and your obj directory grows. It's a useful debugging switch, not necessarily a permanent CI setting.
The headline win is runtime cost. Reflection-based helpers pay for type introspection on every call. Source-generator-based helpers pay nothing extra: the work is done, the code is there, and the JIT treats it like anything else you wrote. For high-throughput paths (serializers, loggers, request dispatchers), the difference between "a few hundred nanoseconds per call" and "tens of microseconds per call" is the whole reason source generators exist.
The structural wins matter as much:
| Concern | Reflection | Source Generator | Hand-Written Code |
|---|---|---|---|
| Runtime cost | High (metadata walk + boxing) | None beyond the call itself | None beyond the call itself |
| Startup cost | High (first-time JIT + reflection caches) | Negligible | Negligible |
| Build-time cost | None | Small per build | None |
| NativeAOT support | Limited; trimmer often strips needed metadata | Full | Full |
| Trimming-friendly | Risky; relies on metadata the trimmer may remove | Yes; generator emits real code | Yes |
| Debuggability | Stack traces show reflection internals | Stack traces show your generated method | Best of all |
| Maintenance burden | Low (one helper handles every type) | Moderate (the generator is a real codebase) | High (handwrite for every type) |
| IDE feedback | Errors at runtime | Errors at build | Errors at build |
The pattern is that generators give you the runtime profile of hand-written code with most of the maintenance profile of a reflection helper. You write the generator once, every marked type gets the boilerplate, and the failure mode is a build error instead of a runtime exception.
The "errors at build" row deserves a moment. A reflection helper that breaks because someone renamed a property produces wrong output at runtime without warning; the only way to find out is to run it. A generator-produced helper either keeps working (because the rename flows through the generator) or fails at compile time (because the generated code references a member that no longer exists). The feedback loop is shorter, and the broken code never ships.
The structure of the generated file deserves a moment. The file is just C#; nothing about it tells the runtime "this was generated". The // <auto-generated/> comment at the top is a convention recognized by various tools (StyleCop, code coverage, the IDE's "find references" feature), but the compiler treats the file like any other.
Consider what changes when the user evolves the source. When the consumer renames Product.Price to Product.Cost:
The next build does this, step by step. The compiler reparses Product.cs and produces a new syntax tree where Cost replaces Price. The semantic model produces a new INamedTypeSymbol for Product whose property list now contains Cost. The generator's transform projects this into its cached value; because the projection contains the property name as a string and the string is different, the cache treats the value as new. The source output step reruns for Product and emits a new Product.ToString.g.cs with Cost in place of Price. The compilation proceeds, and the produced binary has a ToString() that prints Cost correctly.
The key property is that none of this requires the user to do anything beyond the rename. The generator picks up the new shape on the next build, the generated file updates, and the runtime behavior follows. Compare this to a reflection-based helper, where the rename is invisible until the reflection walk runs at runtime; the code still compiles, but PropertyInfo.GetValue now returns a different property than the user might expect. Generators close that gap by moving the discovery from runtime to build time.
This also explains why generators play well with IDE refactorings. A "rename property" command updates Product.Price to Product.Cost in every file in the solution. The generator's output isn't a file in the solution, so the refactoring tool doesn't touch it, but the next build regenerates the file with the new name. The end result is the same as if every reference (including the generated ones) had been renamed manually.
Generators aren't always the best fit. The overhead of writing one (a separate project, a netstandard2.0 target, a non-trivial pipeline) is real, and a small problem rarely justifies it. A few signals point the other way:
ToString() on one class, you write it by hand. The generator is for "every class with [ToString]".dynamic is faster to prototype with.Generators aren't a dynamic replacement either. dynamic defers binding to runtime, which is the opposite of what a generator does. They solve different problems: dynamic is for genuinely unknown shapes at compile time, generators are for known shapes that are tedious to write out.
There's also a category of problem where the generator is technically possible but awkward. Anything that depends on data outside the compilation (a JSON config file, a database schema, a network resource) requires the generator to read that data at build time, which is doable but couples your build to that external system. Teams that go down this path often end up with brittle builds that fail when the external system is unreachable. The standard pattern is to translate the external schema into a code-style description that lives in source control, and then have the generator read that.
The BCL and the broader .NET ecosystem ship a number of source generators. You don't have to write the boilerplate yourself; the generator does it. Each one trades runtime reflection for build-time code emission.
[JsonSerializable(typeof(Order))] to a JsonSerializerContext, and the generator emits efficient, AOT-safe serializers and deserializers. The same code that would otherwise need to reflect over Order's properties at startup is written into the build output. The difference at the call site is small: instead of JsonSerializer.Serialize(order), you write JsonSerializer.Serialize(order, AppContext.Default.Order). The runtime difference is large, because the second form has a real generated method to dispatch to and the first relies on reflection.[LoggerMessage(...)] and the generator produces a high-performance logging method with no boxing of message-template arguments. Replaces the older runtime LoggerMessage.Define pattern. Before this generator, every log call with arguments paid for boxing each value type into object[] for the template substitution. After, the generated method takes the values by their actual types and writes them directly into the logger's LogValues structure.[GeneratedRegex(@"\d+")] and the generator emits a compiled state machine for that pattern at build time, instead of compiling the pattern on first use. The build pays the regex-compilation cost, the runtime pays nothing on the first call, and the result is friendlier to AOT than RegexOptions.Compiled which used to JIT IL at runtime.IRequestHandler<,> at startup, which involved scanning the loaded assemblies and produced startup-time hits and AOT compatibility headaches.[ObservableProperty] and partial methods with [RelayCommand], and the generator emits the property changed notifications and ICommand plumbing that MVVM apps would otherwise write by hand. A property with three lines of user code (a field, a property declaration, and a SetProperty call inside the setter) collapses to one line: [ObservableProperty] private string _name;.Two of these (System.Text.Json and LoggerMessage) are direct replacements for older reflection-based APIs, and the move has been broadly adopted once teams hit AOT or performance ceilings. The rest are net-new conveniences that wouldn't have been possible before incremental generators landed.
The pattern across all of these is the same: a marker attribute on user code ([JsonSerializable], [LoggerMessage], [GeneratedRegex], [ObservableProperty]), a generator that finds the marked declarations through ForAttributeWithMetadataName, and emitted source that turns the declaration into a real, AOT-friendly implementation. Once you've read one generator's pipeline, the rest read like variations on the same theme. The [ToString] example in this lesson uses the same shape that ships in the BCL; the only difference is what BuildToString produces at the end.
A few habits separate generators that ship cleanly from generators that break under load. These are the lessons every author of a working generator picks up the hard way.
Generate the simplest code you can. A generator's job is not to be clever; it's to produce code a human could have written. The simpler the output, the easier the failure mode. If your generated ToString() is twenty lines instead of two, every consumer's stack traces get twenty lines harder to read. Lean toward direct code.
Make the generated file readable. Use the same indentation a human would, put // <auto-generated/> at the top, and emit comments that explain non-obvious choices. The generated file is going to be read by someone, eventually, when something goes wrong. Treat it as a deliverable, not a temporary artifact.
Test the generator like any other library. The Microsoft.CodeAnalysis.Testing packages let you write unit tests that say "given this input, the generator should produce this output". Tests like that catch regressions in the generated code, edge cases (unusual property types, generic outer classes, conflicting names), and pipeline-caching issues. A generator without tests rots, because every C# language version changes the input shape slightly.
Plan for hostile input. A user will eventually mark a struct, an interface, a generic class, or an inherited member with your attribute. The generator should either handle the case correctly or report a diagnostic, never produce broken code. The fastest way to drive users away is for the build to fail with an obscure compiler error from the middle of a generated file.
Keep the pipeline shallow. Two or three steps is enough for most generators. A pipeline with seven cached steps usually has a single complex projection that should be split, or a Combine that's being used to pass data that would be cleaner as a static helper. Complexity in the pipeline is hard to debug because each step's input has to compare correctly to its previous run's input.
These are patterns more than rules. The BCL generators (JSON, regex, logger) follow all of them, and the open-source generators that ship cleanly (MediatR's, MVVM Toolkit's) tend to follow them too. The ones that don't usually end up with bug reports about IDE responsiveness, build errors, or unexpected runtime behavior.
The full incremental pipeline fits in one picture. The three steps from earlier (post-init, syntax provider, source output) feed into each other in a way that maps directly to the cached, incremental nature of the API.
The diagram shows two pipelines that converge. The attribute scaffolding flows straight from RegisterPostInitializationOutput into the final assembly because it doesn't depend on user code. The per-type generated files flow from the syntax provider through the cache and into RegisterSourceOutput. The cache is the part that makes incremental builds cheap: when nothing relevant about Product has changed, the cached INamedTypeSymbol flows straight through and the corresponding output is reused without re-running BuildToString.
The shape repeats across almost every real generator. The post-initialization output emits whatever attribute or interface the generator depends on, the syntax provider finds the user code that uses it, and the source output writes the generated code. The names of the steps change slightly across the BCL generators, but the structure is the same.
The diagram has an implicit ordering. RegisterPostInitializationOutput writes the attribute file. The user code that references the attribute can then compile, which means the attribute is in scope by the time SyntaxProvider.ForAttributeWithMetadataName runs. The syntax provider produces a stream that flows through the cache and into RegisterSourceOutput. If the user adds a new marked class between builds, the cache sees a new value for that class only, and only that class's source output runs. If the user removes a marked class, the cache sees one fewer value, and the corresponding generated file isn't produced on this build (the compiler discards the old one). The compilation that the compiler eventually emits to IL is the union of user source and every generated file from every active generator.
One wrinkle deserves flagging. If your generator emits a file with the same name as another generator's file, the compiler issues an error. The convention is to namespace your file names with your generator's name or namespace prefix, for example ToStringGen.Product.ToString.g.cs instead of plain Product.ToString.g.cs, when there's any chance of overlap. For small private generators this rarely matters; for shared libraries that consumers might combine with other generators, it's a good defensive habit.
A common reason a generator behaves badly is broken cache identity in the pipeline. The compiler caches each step's output by comparing the input to the previous run's input. If the comparison says "same", the cached output flows through unchanged. If it says "different", the step reruns and everything downstream of it also reruns. The art of writing an incremental generator is making sure "same" actually means "same".
Three things break cache identity in practice. The first is using reference types without overriding Equals. A class with default reference equality compares unequal to any other instance, even an instance with identical fields. The cache sees a new reference every time the step runs and concludes the input changed. The fix is to use a record (which generates value equality automatically) or implement Equals and GetHashCode explicitly.
The second is including Roslyn types like INamedTypeSymbol, ISymbol, or Compilation in the cached value. These types are bound to a specific compilation, and the compiler invalidates them between builds. Even if two symbols would describe the same type, they're different instances and compare unequal. The remedy is to extract just the data you need (names, type strings, accessibility flags) before the value reaches RegisterSourceOutput.
The third is putting ImmutableArray<T> or similar collection types into a record without thinking about their equality. ImmutableArray<T>.Equals is reference equality by default, which means two arrays with identical contents compare unequal. The fix is to use SequenceEqual in a custom equality implementation, or to use EquatableArray<T> (a small helper many generator authors copy into their own projects) that wraps an ImmutableArray with structural equality.
The difference, in a sketch:
The trade-off is that you can't carry rich Roslyn data through the pipeline. The translation from INamedTypeSymbol to a small record happens inside the transform callback of ForAttributeWithMetadataName, and from that point on the generator deals only with the projection. This is more work to write, but it's the price of incremental caching that actually works.
A useful test for whether your pipeline values are well-behaved: write a quick unit test that builds the projection from a known type symbol, then builds it again, and asserts the two values are equal. If they aren't, the cache will treat them as different and your generator will rerun everything on every build. Catching the failure in a unit test is far easier than diagnosing it later from IDE responsiveness complaints.
A generator with broken cache identity reruns its source-output step on every keystroke in the IDE. For a small generator this is barely noticeable; for one that emits a thousand files per project, it can make the editor visibly slow. The bug is invisible from the outside; you only catch it by measuring or by reading the generator's pipeline carefully.
Real generators often need to look at more than one thing. A serializer generator might need both the list of [JsonSerializable]-marked types and the project's LangVersion setting; a logging generator might need both the marked partial methods and the project's AssemblyName. The incremental pipeline composes multiple inputs with Combine.
Combine takes two pipelines and produces a single pipeline of pairs. The cache identity rule still applies to the pair, which means both the class info and the assembly name need value equality for the combined pipeline to cache correctly. Records compose nicely here because the compiler's generated Equals walks every field.
This is the same composition pattern you'd see in the official LoggerMessage and Regex generators inside the BCL. They often Combine candidate methods with the compilation's language version or with a flag that says "is nullable enabled in this project", because the emitted code differs depending on those settings. You don't have to use Combine for every generator, but knowing it's there saves you from the temptation to smuggle extra data through closures or static fields, both of which break caching in ways that are hard to debug later.
C# has had multiple code-generation stories over the years, and source generators replace most of them. Knowing the older approaches helps explain why generators look the way they do.
T4 templates (Text Template Transformation Toolkit) were the previous mainstream answer for "generate code at build time". A .tt file mixed C# fragments with text, and an external tool ran the template to produce a .cs file that lived in source control. T4 worked, but it had hard edges: the templates ran in a separate process, the tooling was Windows-centric and tied to Visual Studio, the generated files were checked into the repo and had to be regenerated by hand when inputs changed, and incremental rebuilds were a mess. Source generators address every one of those issues. The generator runs inside the compiler, on every platform dotnet runs on, and the generated files live in obj/ (never in source control), so the build is the only thing that ever writes them.
`Reflection.Emit` is the runtime-side alternative. You build a dynamic assembly in memory, emit IL instructions into methods, and load the result. This is what older serializers (DataContractSerializer, the Compiled mode of XmlSerializer, some versions of Newtonsoft.Json) did internally: at first call, they'd reflect over the type, generate optimized IL, and cache it. The runtime cost of subsequent calls was good, but the first call paid a heavy price, the generated IL was invisible to debuggers, the trimmer couldn't see it, and AOT publication couldn't include it. Source generators move the same work to build time, where the trimmer and the AOT compiler can analyze the produced code normally.
Expression trees are a third option, building a Expression<Func<T, U>>, compiling it to a delegate at runtime, and caching the delegate. They're cleaner than Reflection.Emit because you build trees with normal-looking C# instead of raw IL, and the JIT understands the compiled output. They still have the same AOT and trimming issues, though, because the compiled delegate is JIT-only. Generators are the AOT-safe answer for most expression-tree use cases.
The general pattern across these alternatives is the same: each one trades off where the work happens (build vs first call vs every call), what tooling supports it (cross-platform vs Windows-only, AOT-friendly vs JIT-only), and how visible the output is (in source control vs in memory vs in obj/). Source generators are the better fit on every axis for problems whose types are known at build time. They don't compete on problems whose types are only known at runtime, which is where reflection and its derivatives still have a role.
Generators run inside the compiler process, which makes them a little harder to debug than ordinary code. A few practical tips:
EmitCompilerGeneratedFiles to true and read the actual output. Half of all generator bugs become obvious the moment you can read the generated file.ctx.ReportDiagnostic) when the input is malformed, so users get a clear build message instead of a confusing downstream error.dotnet build-server shutdown works).System.Diagnostics.Debugger.Launch() inside the generator to pop the attach-debugger dialog, then step through the generator with your IDE attached to the compiler process. This is rarely necessary for small generators but invaluable for big ones.These are workflow items, not features of the language. The generator is code; you debug it like code, with the added wrinkle that its host is a compiler.
A generator that emits "wrong" code without warning when the user makes a mistake leaves people guessing. The better pattern is to report a proper diagnostic, the same kind of squiggly-line message the compiler itself produces. The incremental generator API exposes this through SourceProductionContext.ReportDiagnostic.
The shape is familiar from compiler errors: an ID (which users can suppress with #pragma warning disable TOSTR001 if they really need to), a human-readable title, a parameterized message, and a severity. The location is what makes the diagnostic clickable in the IDE; passing classSymbol.Locations.FirstOrDefault() points the squiggly line at the user's declaration, not the generator's code.
Severity levels matter too. Use DiagnosticSeverity.Error to fail the build for problems that would otherwise produce invalid generated code. DiagnosticSeverity.Warning lets the build continue but flags the issue in the IDE, which fits cases like "this is suboptimal but still works". DiagnosticSeverity.Info and DiagnosticSeverity.Hidden are quieter still and rarely worth using from a generator; they're more common in analyzers.
Diagnostics turn vague failures into precise messages. Without them, a user who forgets partial sees a CS0260 from the compiler ("Missing partial modifier...") and has to figure out that it came from a generator they may not even know is running. With a diagnostic, the user sees "Class 'Product' is marked with [ToString] but is not declared partial" and knows exactly what to do. The investment is small, the user experience improvement is large, and it's how every well-behaved generator in the BCL handles invalid input.
The full sequence a single build goes through with your generator in the pipeline is conceptually simple.
The compiler parses your source, runs every registered generator, takes the files those generators produced, merges them into the same compilation, and then proceeds with the rest of the build as if you'd written everything by hand. If a generator reports a diagnostic, that diagnostic shows up alongside any compiler-issued ones. If a generated file has a syntax error, that error also shows up; from the compiler's perspective, generator output is just more source code, no more privileged than what you wrote.
The IDE follows a similar path but with caching turned way up. Every keystroke produces a new logical compilation; the incremental pipeline reuses the previous step's output everywhere the inputs match. When you edit a class that has no [ToString] attribute, the generator's syntax provider sees that nothing changed for it and the source output step doesn't rerun. When you add [ToString] to a new class, the syntax provider produces a new candidate, that candidate flows through the cache as a new value, and the source output step runs for just that class. None of the previously generated files are touched.
This is the practical reason IIncrementalGenerator exists. Without caching, the IDE would have to rerun every generator on every keystroke, and a project with a handful of generators (the JSON one, the regex one, your own) would grind to a halt. With caching, only the work that actually depends on the change happens, and the editor stays responsive.
To bring everything together, walk through what happens end-to-end when a developer adds [ToString] to a new Order class for the first time.
The developer edits Order.cs, adds using ToStringGen; at the top, and writes [ToString] public partial class Order { public int Id { get; init; } public decimal Total { get; init; } }. They save the file and run dotnet build.
The compiler picks up the change. It parses every source file in the project, including the new version of Order.cs. It runs the registered generators. Your ToStringGenerator is one of them; its Initialize method has already wired up the pipeline during compiler startup. RegisterPostInitializationOutput produces ToStringAttribute.g.cs (which has been the same for every build, so the cache reuses the cached file). SyntaxProvider.ForAttributeWithMetadataName walks the compilation looking for classes marked with ToStringGen.ToStringAttribute. It finds Product, Customer, and the new Order. The first two flow through the cache unchanged because their syntax hasn't changed since the previous build. Order is new, so the cache produces a new value for it and pushes that value into RegisterSourceOutput. The source-output step runs BuildToString(orderSymbol), which produces a small Order.ToString.g.cs containing the partial class Order { public override string ToString() => $"Order { Id = {Id}, Total = {Total} }"; } declaration. The compiler adds this file to the compilation, type-checks the result, and emits the final assembly.
The developer hits F5 and runs the program. They see Order { Id = 1, Total = 99.50 } printed without ever having written ToString() themselves. The generator didn't run when the program executed; the runtime called the ordinary method the build had produced. The only sign the generator was ever involved is the file under obj/generated/..., which the developer can open to read what the build did on their behalf.
That's the full loop. The patterns inside (partial, [Generator], the three pipeline steps, the cache, the diagnostic mechanism) all exist to make this loop reliable, incremental, and predictable across editor sessions and build environments. The benefit is that an entire category of boilerplate moves out of the user's editor and into a generator that runs once during the build, with the same runtime cost as code the user could have written by hand.
9 quizzes