You paste a paragraph out of a PDF, a Word document or a web table and the spacing is wrong: gaps too wide, ragged indentation, four blank lines where you wanted one. So you reach for find-and-replace, swap two spaces for one, and it reports two replacements - while the gap you were actually staring at is still there. This guide explains why, what the characters in your text really are, and the order the cleanup has to run in.
Why the gap will not close
Almost every piece of advice about extra spaces assumes the gap is made of spaces. Often it is not. The space you type, U+0020, is one of roughly twenty characters that render as a horizontal gap, and word processors, PDF generators and web pages use the others routinely. On the clipboard they are indistinguishable by eye.
The most common by far is the no-break space, U+00A0. It is what a browser produces from the HTML entity , what Word inserts to keep a number with its unit, and what a PDF viewer hands over from justified text. It looks exactly like a space and is a different character, so a search for two spaces does not find a pair of them, and a tool built on that search reports your text is clean. Hence the conclusion that there is some invisible problem you cannot fix.
A worked example, character by character
Take an invoice reference copied from a web table. On screen it reads Invoice No. 2026-0417, with two wide gaps and stray indentation. On the clipboard it is 27 characters:
- Two ordinary spaces at the start.
- Invoice - 7 characters.
- Two no-break spaces (U+00A0), which is the wide gap you can see.
- No. - 3 characters.
- Two tab characters, which is the second wide gap.
- 2026-0417 - 9 characters.
- Two ordinary spaces at the end.
Find-and-replace on two spaces reports two matches: the pair at the start and the pair at the end. The two gaps in the middle - the ones you were trying to fix - are untouched, because neither is made of spaces. Paste the same text into the Whitespace Remover and it names what it found: 2 no-break spaces, 2 tabs. The output is Invoice No. 2026-0417, 21 characters, 6 removed and 4 gaps fixed.
The characters that look like a space but are not
These all render as a gap, so each is safe to turn into an ordinary space:
- No-break space (U+00A0) - by far the most common, from HTML, Word and PDFs.
- Narrow no-break space (U+202F) - a thousands separator in French and Russian number formats.
- Figure space (U+2007) - the width of a digit, used to align numbers in columns.
- Em space (U+2003), en space (U+2002), thin space (U+2009), hair space (U+200A) - typographic spacing from design tools and typeset PDFs.
- Ideographic space (U+3000) - the full-width space used in Chinese, Japanese and Korean text.
A tab, U+0009, belongs here for a different reason: its width is whatever the thing displaying it decides, so a tab in a form field or a CSV cell is a gap of unpredictable size. Collapsing runs of spaces and tabs together is what makes pasted text behave.
The invisible ones
A second group has no width at all, so nothing looks wrong - but they count toward character limits and break exact-match comparisons. It is why a code, a password or a search term that looks identical refuses to match.
- Zero-width space (U+200B) - a line-break hint, commonly pasted from web content.
- Byte-order mark (U+FEFF) - left at the start of a badly saved UTF-8 file; the classic broken first CSV header.
- Soft hyphen (U+00AD) - inserted by PDFs where a word was split across lines, so it arrives with a hyphen hiding inside it.
- Left-to-right and right-to-left marks (U+200E, U+200F) - direction hints from mixed-script text.
These should be deleted, not converted to a space. A zero-width space inside a word is a line-break hint, not a gap between words, so replacing it with a space splits the word in two - the difference between cleaning text and damaging it.
The two you must not delete
The zero-width joiner (U+200D) and non-joiner (U+200C) are invisible too, so tools advertising invisible-character removal sweep them up with the rest. They carry meaning. A family emoji is several emoji bound together by joiners, so stripping them turns one picture into three. In Devanagari and Perso-Arabic they decide whether a conjunct renders as a ligature or a half-form, so removing them corrupts Hindi, Arabic and Persian - every letter still there, every word spelled wrong. A cleaner should leave these alone by default and make removing them a deliberate, separate choice, which is how the Whitespace Remover treats them.
Trailing spaces and blank lines
Trailing spaces at the ends of lines are invisible and survive every paste. They make a list look ragged in a form field, they break a spreadsheet lookup, and leading spaces do the same at the other end.
Blank lines need a decision rather than a rule. One between paragraphs is a paragraph break and should stay; a run of four is an artefact of the layout the text came from. Collapsing a run to one keeps prose readable, removing them all closes up a list, and blank lines at the very start and end should just go - an empty first line is not a paragraph break.
Doing it in a spreadsheet or an editor
In Excel or Google Sheets, TRIM only half works: it strips leading and trailing spaces and collapses internal runs, but it does not touch U+00A0, so the cell that looked wrong still does. CLEAN does not help either - it only removes ASCII control characters below code 32. The working formula substitutes first: TRIM(SUBSTITUTE(A1, CHAR(160), CHAR(32))). Each zero-width character needs its own SUBSTITUTE on top.
In an editor with regex search, replacing \s+ with one space gets further than you might expect, because JavaScript's \s does include the no-break space and the typographic spaces. But it also matches newlines, so it flattens paragraphs into one line, and it misses the zero-width space entirely. \s is not the set you assume in either direction: it excludes U+200B, which people paste constantly, and includes U+FEFF, which nobody types.
The order the steps have to run in
The steps interfere with each other. Normalise line endings first: a file using carriage returns, or a U+2028 pasted out of a JavaScript string, is a single line to every per-line operation, so nothing else will fire. Delete the zero-width characters next, before collapsing gaps - a space, then a zero-width space, then a space is not a run of whitespace while that character sits in the middle, so collapsing first leaves two spaces behind. Convert the exotic spaces third, so the collapse sees one kind of character. Only then collapse runs, trim line ends and handle blank lines.
Get the order wrong and the text looks cleaned while still holding what you were removing - the worst outcome, because you stop looking. Run it in that order and the result is stable: cleaning it again changes nothing.
Frequently asked questions
- Why does find-and-replace not remove my double spaces?
- Because the gap is probably not made of space characters. Text from Word, PDFs, web pages and spreadsheets is full of no-break spaces (U+00A0) and typographic spaces like the em space and figure space, which render exactly like a space but are different characters. Searching for a space does not find them, so the tool truthfully reports no matches while the gap you can see stays put. The fix is to convert those characters to ordinary spaces first and collapse afterwards. Any cleaner worth using will also tell you which characters it found by name, because knowing your text contains four no-break spaces explains the CSV column that would not parse, the lookup that returned nothing, and the reference that would not match.
- How do I remove all spaces from text rather than just the extra ones?
- Choose the option that removes all spaces rather than collapsing them. Every space and tab is deleted, including single spaces between words, so h e l l o becomes hello. Line breaks are kept, so each line of a list keeps its own line and just loses its internal spacing. Use the default collapse setting for ordinary prose, where you want runs of whitespace reduced to a single space and the text left readable. Removing every space is mainly useful for cleaning up a code, a serial number or a value that was pasted with spacing inside it.
- Is it safe to remove invisible characters from text with emoji or Hindi?
- Only if the tool separates the two kinds. Characters like the zero-width space, byte-order mark and soft hyphen carry no meaning in plain text and can be deleted safely. The zero-width joiner (U+200D) and non-joiner (U+200C) are different: a family emoji is several emoji bound together by joiners, so deleting them turns one picture into three, and in Devanagari and Perso-Arabic they decide whether a conjunct renders as a ligature or a half-form, so removing them corrupts Hindi, Arabic and Persian text without any visible error. A cleaner should leave those two alone by default and make removing them a separate, clearly labelled choice.