Developer & encoding

How to Escape a String for JSON, HTML, XML, CSV and SQL

How to escape a string properly: what changes per format, the ordering bug in nearly every hand-rolled escaper, the characters XML cannot hold at all, and why escaping is not a security control.

6 min readUpdated Aug 30, 2026

Sooner or later a piece of text breaks something. A customer called O'Brien ends a SQL query halfway through. A product description with an ampersand stops an XML feed parsing. A company name containing a comma splits into two spreadsheet columns. The fix in every case is escaping - rewriting the characters the destination treats as punctuation so they arrive as plain text. This guide covers what escaping changes, the mistake nearly every hand-rolled escaper makes, and where escaping stops being enough. The String Escaper does all eight formats below in your browser.

Escaped for what?

The most useful thing to know is that there is no such thing as an escaped string. Escaping is always escaping for something, and the rules do not overlap. Take one apostrophe. In JSON nothing happens to it. In an HTML attribute it becomes ', in XML ', in SQL two apostrophes. On a command line it turns into a four-character sequence that closes the quoted string, adds an escaped quote and reopens it. Five destinations, five answers, and applying the wrong one is not a partial fix - it is a new bug.

So the first question is never how to escape but where the text is going. That is why the tool asks for a target format first, and why each one names the exact position its output is valid in: the inside of a JSON string, one RFC 4180 field, one shell argument.

One sentence, eight formats

Here is a line of ordinary marketing copy - an ampersand, an apostrophe, two double quotes and a comma:

  • Tom & Jerry's "Best of", 50% off

Paste it into the String Escaper and switch the target format:

  • JSON: Tom & Jerry's \"Best of\", 50% off
  • JavaScript (double quotes): Tom & Jerry's \"Best of\", 50% off
  • HTML: Tom & Jerry's "Best of", 50% off
  • XML: Tom & Jerry's "Best of", 50% off
  • CSV: "Tom & Jerry's ""Best of"", 50% off"
  • SQL: Tom & Jerry''s "Best of", 50% off
  • Shell: 'Tom & Jerry'\''s "Best of", 50% off'

Five different treatments of the same 32 characters. JSON touches only the double quotes. HTML and XML rewrite four characters each but disagree on the apostrophe, because ' was an XML entity that HTML did not define until HTML5 - the numeric ' is the form every browser has always understood. CSV leaves the interior alone but wraps the whole field in quotes because of the comma, doubling the interior ones. SQL changes a single character. The shell wraps the lot in single quotes, where nothing expands at all, and performs its close-escape-reopen dance around the apostrophe.

The regular expression format is the odd one out, because this sentence has no metacharacters. Give it one that does - price: $9.99 (50% off) - and it returns price: \$9\.99 \(50% off\), which matches that text and only that text.

The bug in almost every hand-rolled escaper

Escaping HTML looks like three lines of code, and the obvious three lines are wrong:

  • text.replace('<', '&lt;').replace('>', '&gt;').replace('&', '&amp;')

Run <b> through that and you get &amp;lt;b&amp;gt; - a visible mess on the page instead of a bold tag. The first replace introduced ampersands and the last escaped them again. The ampersand has to go first, because it is the character every other escape is built out of. The rule is obvious once you know it; it survives in so much code because the output is still valid HTML, just wrong, so nothing crashes until a customer sees it.

Building the output one character at a time, rather than by a chain of replacements, makes the bug structurally impossible - which is how this tool works. It also handles what a regex-based escaper gets wrong quietly: astral characters, and unpaired surrogates left when a string was truncated at the wrong byte.

The characters that cannot be escaped at all

Most escaping questions have an answer. A few do not, and XML has the sharpest example. XML 1.0 permits exactly three control characters - tab, line feed and carriage return - and there is no escape for the rest. Writing &#1; does not produce character one; it is itself a parse error. A document holding a stray control character, common in text pulled from a legacy database, cannot be made into valid XML 1.0 without removing something.

A tool that just swaps the five predefined entities hands you that document and says nothing; you find out when the parser at the other end rejects it. The String Escaper names the code points instead and offers to strip them. Whitespace holds a related trap: a parser normalises a literal carriage return to a line feed, and inside an attribute value turns tabs and newlines into plain spaces. If those matter they must be written as &#13;, &#9; and &#10;, which is what the attribute-value setting does.

Escaping is not a security feature

This is worth saying plainly, because escaping is often reached for as a defence. Escaping a SQL value protects it only where it already sits inside quotes. It does nothing for:

  • A numeric column, where the value is not quoted at all.
  • A table or column name, which cannot be a bound parameter in most drivers.
  • An ORDER BY direction, a LIMIT, or any other fragment of the statement itself.

It also has to match the exact dialect. MySQL and MariaDB read a backslash as an escape character and standard SQL does not, so the same escaped string means two different things depending on whether NO_BACKSLASH_ESCAPES is set. A parameterised query avoids all of this by sending the statement and the value separately, leaving no parsing step to subvert. Escape by hand only when you are genuinely writing a literal, in a migration or a one-off report. The same caution applies to HTML: escaping makes text safe as element content and inside a quoted attribute, but not inside a script or style block, and not in a URL attribute, where a javascript: scheme needs a different check.

The two that cannot be undone

Six of the eight formats reverse cleanly, so the tool offers an Unescape direction for them - useful when a log line arrives as a JSON string, or a scraped page is full of &eacute; and &#151;. Two do not. A regular expression cannot be reliably unescaped, because a backslash before a letter usually does not mean that letter: \d is any digit, \b a word boundary, \w a word character. Shell quoting has the opposite problem - the same argument can be written many equally valid ways. Rather than guess, the tool offers those two in one direction and says why.

It all runs in your browser

The text you paste into an escaper is often what you would least like to send anywhere: a connection string, an API key, a customer record, a password you are quoting for a command line. Every rule here runs in JavaScript inside your tab, nothing is uploaded, and the page works with the network off. For a whole document rather than one value, the JSON Formatter and XML Formatter handle those; URL Encoder covers percent-encoding, and Base64 Encoder the other common transport encoding.

Frequently asked questions

What is the difference between escaping and encoding?
Escaping rewrites only the few characters that would otherwise be read as punctuation by a particular parser, leaving the rest of the text readable - a JSON string with one escaped quote in it still looks like the original sentence. Encoding transforms the whole value into a different alphabet, usually so it can survive a channel that cannot carry arbitrary bytes: Base64 turns any data into 64 safe characters, and percent-encoding rewrites a URL component. Escaping is about a parser's grammar; encoding is about a transport. They are often needed together, and in that case the order matters, because escaping already-encoded text and encoding already-escaped text give different results.
Do I still need to escape if I use a template engine or an ORM?
Usually not, and you should prefer that. Modern template engines escape interpolated values for HTML by default, and an ORM or any driver that supports bound parameters keeps SQL values out of the statement entirely, which is stronger than escaping because there is no parsing step to attack. Hand escaping is for the gaps: a string literal you are pasting into source code, a value going into a hand-written migration, one field of a CSV export, a command line you are assembling, or a regular expression that has to match user text literally. If a library will do it for you in the right context, let it.
Is my text uploaded anywhere?
No. Every escaping rule runs in JavaScript inside your browser tab, so a connection string, an API key, a customer record or a password you are quoting for a command line never leaves your device. Nothing is stored, and the page keeps working with the network switched off.