Projection reshapes each element of a sequence into a new shape, and grouping reorganizes a flat sequence into keyed buckets. Together they cover the "what does each row look like in the output?" and "how do I bucket these by category?" questions that show up in almost every LINQ query. This lesson walks through Select, SelectMany, Zip, GroupBy, and ToLookup, with a focus on what each one returns and when to use it.
Select for Element-by-Element ProjectionSelect is the standard element-by-element projection operator. You give it a function from T to TResult, and it produces a new sequence where every element is the result of applying that function. The shape of the output is whatever you return from the selector.
The input is a sequence of tuples. The selector p => p.Name returns a string, so names is an IEnumerable<string>. Select doesn't filter, it doesn't reorder, and it doesn't change the count. One element in, one element out, always in the same order.
You can project into a completely different type. Here's a decimal projection that turns prices into prices with tax added.
Each decimal becomes another decimal. The output stays the same length as the input. The pattern is the same whether you project to a primitive, a tuple, an anonymous type, or a named class.
Select with an IndexSelect has a second overload that passes both the element and its zero-based index into the selector. The index is useful when you need positional information, like numbering rows for display.
The index always reflects the position in the input sequence at the moment Select runs over it. If a Where runs before this Select, the index here is the position after filtering, not the original position in the source. Order the operators carefully when both are involved.
When you only need a temporary shape inside a single method, anonymous types are the lightest option. The new { ... } syntax creates a compiler-generated, read-only type with the listed properties.
The anonymous type has two properties, Name and Price. The compiler generates a class internally, with read-only auto-properties, a constructor, structural Equals, GetHashCode, and ToString. You can't name the type yourself, so you can't return it from a method signature (the return type would have to be object or dynamic, both of which throw away the type information).
Property names can be inferred from the source expression (p.Name produces a property called Name) or named explicitly with Property = expression syntax.
Two anonymous types with the same property names and types in the same order are considered the same type by the compiler, which is why this projection produces a homogeneous sequence rather than a sequence of distinct ad-hoc types.
Anonymous types allocate one object per element on the heap. For a million-element sequence projected into anonymous types just to print them, that's a million allocations. When the projection only feeds a single immediate consumer (a foreach that reads two fields), prefer pulling the fields directly inside the loop and skipping the projection entirely.
Anonymous types fit when the projection stays inside one method. The moment you want to return the projection, pass it to another method, or store it in a field, you need a named type.
| Situation | Use |
|---|---|
| Temporary shape inside one method | Anonymous type (new { ... }) |
| Return value, parameter, or field | Named record or class |
| Crossing project / assembly boundaries | Named record or DTO class |
| Public API surface | Named record or class with explicit shape |
A record is usually the most direct replacement. Records give you value-based equality, concise syntax, and a readable name in stack traces and signatures.
The method signature IEnumerable<ProductSummary> says exactly what shape callers get back. An anonymous type couldn't fit there, because the type has no name to write.
A rule of thumb: if the projection escapes the method (returned, stored, exposed), use a record. If it dies inside the method, an anonymous type is fine.
SelectMany for FlatteningSelect produces one output per input. SelectMany produces zero, one, or many outputs per input, then concatenates them all into one sequence. The classic use case is flattening a nested collection.
The input has three orders. The output has six lines, because SelectMany took each order's Lines list and concatenated them all into one flat sequence. If you used Select(o => o.Lines) instead, you'd get an IEnumerable<List<OrderLine>>, a sequence of three lists, not a flat sequence of six lines.
SelectMany opens each box and tips its contents into one pile. The order is preserved: lines from order 1 come first, then order 2, then order 3.
If you've worked with SQL, this is analogous to a CROSS APPLY against a child table: each parent row is joined to its child rows and the child rows become the result. The full SQL semantics are out of scope here, but the intuition transfers.
SelectMany with a Result SelectorThe two-argument form of SelectMany lets you keep a reference to the parent while iterating the children. The first argument is the collection selector (the children to flatten), and the second is a result selector that takes (parent, child) and returns the output shape.
Each output row carries both the parent's Id and Customer and the child's Product and Quantity. This is the shape-preserving flatten-and-join pattern. Without the result selector, you'd lose the parent and need another lookup later. With it, you assemble the final row in one pass.
Zip for Combining Two SequencesZip takes two sequences and walks them in parallel, pairing elements at the same index. The output stops at the shorter of the two inputs.
The single-argument overload returns an IEnumerable<(TFirst, TSecond)>, a sequence of tuples. The output has three pairs because quantities runs out first. The fourth product (Monitor) is silently dropped.
The result-selector overload lets you produce a custom shape per pair.
Use Zip when you have two related sequences that you know are positionally aligned, like names and prices coming from a CSV with matching columns. If positional alignment is fragile (the lists could drift out of sync), prefer joining on a key with a proper join operator.
Zip is lazy and pulls one element from each source per output, so it doesn't materialize either input. The truncation to the shorter sequence is silent. If you need to detect mismatched lengths, check Count on both sides before zipping or use a separate validation step.
GroupBy for Keyed BucketsGroupBy takes a flat sequence and a key selector, then reorganizes the elements into groups, one group per distinct key. The output is an IEnumerable<IGrouping<TKey, TElement>>.
The result is a sequence of groups, where each group has a Key property and is itself an IEnumerable<Product>. The outer foreach walks the groups; the inner foreach walks the elements inside each group.
The shape of the reshape looks like this:
The flat input on the left is reorganized into three buckets on the right, one per distinct category. Note that the relative order of products within a group matches their order in the source: Headphones, Speakers, Earbuds, in that order, because that's the order they appeared in the input.
IGrouping<TKey, TElement> IsIGrouping<TKey, TElement> is just IEnumerable<TElement> plus a Key property. That's the whole interface.
That means anywhere you can use IEnumerable<T>, you can use a group, including in another LINQ operator chain or another foreach. The Key property tells you which bucket you're in; the iteration gives you the elements that landed there.
GroupBy with an Element SelectorThe two-argument overload lets you project each element as it goes into its group, instead of storing the full original element. This is useful when the group only needs one field.
The element selector p => p.Name says "store only the name in each group." Now each group is an IGrouping<string, string> rather than IGrouping<string, Product>. The grouping shape is the same; only the element type changed.
GroupBy with a Result SelectorThe three-argument overload skips the intermediate IGrouping and lets you produce one output object per group directly. Use this form when you want a flat summary, one row per key.
The result selector receives the key and the group's items, and returns whatever shape you want. Here we project to an anonymous summary with a count and a minimum price per category. (The aggregate operators Count and Min are covered in depth in the Aggregation chapter; using them inside a group like this is the natural pairing.)
GroupBy is lazy until iterated, but the first iteration walks the entire source once and builds the lookup table in memory. You can't get the first group's items before reading every element of the source, because any element could belong to any group. For large inputs that you only want to peek at, consider whether you need grouping or just need Where.
ToLookup for Immediate, Indexable GroupsToLookup looks similar to GroupBy at first glance, but the two have important differences. ToLookup runs immediately (it materializes the lookup the moment you call it), returns an ILookup<TKey, TElement>, and is indexable like a dictionary.
Two details. First, lookup["Audio"] returns an IEnumerable<Product> directly, no enumeration of groups required. Second, asking for a key that doesn't exist returns an empty sequence, not a KeyNotFoundException. That last detail is the most common reason to pick ToLookup over ToDictionary: a dictionary throws on a missing key; a lookup returns empty.
The table that captures when to use each:
| Operator | When it runs | Return type | Indexable | Missing key |
|---|---|---|---|---|
GroupBy | Deferred (lazy) | IEnumerable<IGrouping<TKey, TElement>> | No, must iterate | N/A |
ToLookup | Immediate | ILookup<TKey, TElement> | Yes, lookup[key] | Returns empty sequence |
ToDictionary | Immediate | IDictionary<TKey, TElement> | Yes, dict[key] | Throws KeyNotFoundException |
Use GroupBy when you want to walk through groups once, possibly chaining more LINQ on top. Use ToLookup when you'll be doing repeated lookups by key, like an in-memory join or a category index. Use ToDictionary when each key maps to exactly one value and missing keys should be an error.
Keyboard wasn't reviewed at all. The lookup returns an empty sequence rather than throwing, which makes downstream code cleaner: no try/catch, no defensive ContainsKey, just call Count() and get zero.
A small end-to-end example combining SelectMany, GroupBy, and a result selector. We have a list of orders, each with multiple lines, each line referencing a product with a category. We want one row per category showing how many units were ordered.
The flow: SelectMany flattens the three orders into a single sequence of six order lines. GroupBy then re-groups those lines by the product's category, and the result selector totals up the Quantity per group. Audio gets 1 + 3 + 2 = 6 units (headphones, earbuds, headphones). Accessories gets 2 + 1 + 1 = 4 units (mouse, keyboard, mouse).
This is the everyday shape of analytics-style LINQ: flatten, then group, then summarize. The full aggregation toolbox lives in the _Aggregation Operators_ lesson, but Sum inside a group selector is natural enough that it shows up here too.
10 quizzes