The obvious idea is to join the strings with a delimiter, like "neet,code,love,you", and split on commas to decode. That breaks the moment a string contains a comma. The constraint says strings can hold any of the 256 ASCII characters, so no single character is safe to reserve as a separator.
The real challenge is making the encoding unambiguous. We need to combine strings into one string and always reconstruct the exact original list, regardless of what the input strings contain. That includes empty strings, strings with special characters, and strings that look like parts of our encoding format.
The scheme has to distinguish "structure" (separators, lengths) from "content" in a way the data can never imitate. Two standard ideas achieve this: escape the delimiter so content can never produce a real one, or record each string's length so the decoder reads by count instead of searching for a separator.
0 <= strs.length <= 200 → The list can be empty, so decoding an empty encoded string must return an empty list.0 <= strs[i].length <= 200 → Strings can be empty. An empty string must round-trip correctly, both in the middle of the list and as the sole element.strs[i] contains any possible characters out of 256 valid ASCII characters → This rules out any plain-delimiter scheme. No character is safe to reserve as a separator, which forces us toward escaping or length prefixes.If we want to keep using a delimiter, we have to make sure a string can never produce that delimiter on its own. Escaping does exactly that: pick a delimiter, and whenever the delimiter character appears inside a string, rewrite it as a longer sequence that the decoder knows to collapse back.
CSV files and JSON solve the same problem this way. CSV doubles a quote inside a quoted field, JSON writes \" for a literal quote. Here we use # as the building block: a literal # inside a string becomes ##, and the two-character sequence #, marks the boundary between strings. Because every real # in the content is doubled, a lone # followed by , can only be a separator, never part of the data.
The encoding round-trips correctly, but the decode logic has to track whether each # is the start of an escaped pair or a separator, which makes it the more error-prone of the schemes here.
# with ##, then append #, as the terminator. Every string, including the empty string, contributes its terminator, so the encoding ends in #,.#, look at the next character. ## means a literal # belongs to the current string; #, means the current string is complete. Any other character is copied as-is.Reading # pairs left to right is what keeps decoding unambiguous. After encoding, every literal # sits at the start of an even-length run of #, so when the scanner reaches a # it always consumes two characters before re-checking. A separator #, therefore can never be mistaken for the middle of an escaped pair.
The escaping scheme is correct but the decode loop has three cases to keep straight, and an off-by-one in the #-pair handling silently corrupts the output. The next approach removes escaping entirely by recording how many characters each string occupies.
Instead of escaping special characters, prepend each string with its length. Once the decoder knows how many characters belong to the next string, it reads exactly that many and moves on, no separator needed between strings.
The one ambiguity is where the length ends. Encoding ["hi","bye"] as "2hi3bye" leaves the decoder unable to tell the count 2 from the start of the content. A single separator between the length and the content resolves it: write "2#hi" and "3#bye". The decoder reads digits until the #, parses the count, then reads that many content characters.
This is safe even when a string contains # or digits, because the decoder only searches for # while it is reading a length. Once it has the count, it switches to reading by count and never inspects the content for separators.
Length-prefixed framing is a standard technique. TCP frames messages by byte count, HTTP uses the Content-Length header, and Protocol Buffers use length-prefixed fields. The principle is the same: state how much to read before reading it.
The case to verify is a string that contains the separator. Encoding ["a#b"] produces "3#a#b". Decoding reads digits until the first #, giving length 3, then reads exactly 3 content characters: a, #, b. The second # falls inside the content window and is never tested as a separator, so the string survives intact.
#, then append the string.# to get the length, then read exactly that many characters as the next string. Repeat until the encoded string is fully consumed.This handles every edge case and is the standard solution to the problem. Decoding still scans forward to find each #, though. The next approach removes that scan by giving every length a fixed width.
Give every length a fixed width and the separator scan disappears. The maximum string length is 200, which fits in three digits, so a 4-digit zero-padded header (with one digit to spare) holds any length the constraints allow. The decoder reads exactly 4 characters for the length, then reads that many characters for the string.
No separator character is needed. The format is self-describing: the first 4 characters are the length, the next N characters are the string, the following 4 characters are the next length, and so on. Decoding becomes pure index arithmetic with no inner search.