Back to blog
Developer Tools Data cleaning Regex

How to Extract Emails, Phones, URLs & IPs from Text (Split & Regex)

Published: Updated: Reading time: about 6 min
Share:

Server logs, export files, group chats and customer notes are full of useful data — but it's buried in prose. Manually copy-pasting emails or phone numbers out of a 10,000-line file is slow and error-prone. This guide shows how to extract emails, phone numbers, URLs, IP addresses and IDs with preset or custom regular expressions, split text into columns, and deduplicate the results — all in your browser.

When You Need a Text Extractor

ScenarioWhat you pull out
Support tickets & chatsEmail addresses, phone numbers, order IDs
Server / app logsIP addresses, URLs, dates, error codes
CSV & spreadsheet exportsColumns split by delimiter, deduplicated rows
Market researchCurrencies, hashtags, mentions, hex colors

Extract with Preset Patterns

Most extraction tasks are covered by a handful of well-known shapes. The text extractor ships with presets for the most common ones:

  • Email — the classic name@domain.tld shape.
  • URL — http(s)://… links.
  • IPv4 / IPv6 — dotted-quad and colon addresses.
  • Phone — a lenient international pattern (with/without country code, dashes, parentheses).
  • Number — integers and decimals, skipping IP segments.
  • ID — hyphenated codes like INV-2026-0042.
  • Date, Hex color, Hashtag, @mention, Currency — niche but common.

Presets are intentionally lenient: they find candidates fast, and you refine precision with a custom regex when needed.

Custom Regex for Precision

Presets cover 90% of cases; the last 10% needs your own rule. Type a pattern like \b[A-Z]{2}\d{4}\b to match exactly 2 uppercase letters followed by 4 digits, and toggle the g (global), i (case-insensitive) and m (multiline) flags. The tool validates your pattern and reports a friendly error if it's invalid.

Tip: wrap boundaries (\b or lookarounds) in custom patterns. Without them, a pattern like \d{11} will also match the middle 11 digits of a longer number.

Deduplicate Results

Logs and exports repeat themselves constantly. Tick "Remove duplicates" and the output collapses to unique values, with the stats line showing exactly how many duplicates were dropped: Found: 342 → 58 unique · 284 duplicate(s) removed.

Split Text into Columns

When your data is already structured but separated by delimiters (CSV from a spreadsheet, TSV from a database export), the Split mode is faster than regex:

  1. Pick a delimiter — comma, tab, space, semicolon, pipe, newline, or a custom one.
  2. Review the table — rows render as a live table for a quick sanity check.
  3. Advanced options — remove duplicate rows, sort by any column (numeric or text, ascending/descending), or transpose rows ↔ columns.
  4. Copy as CSV — quoted and escaped correctly, ready for Excel or Google Sheets.

Frequently Asked Questions

Why doesn't my preset match everything I expected?

Presets are lenient but not perfect — phone formats vary wildly by country, and IDs are often project-specific. If a preset misses cases, copy it into the custom regex field and adjust, or write your own pattern.

How do I extract only unique values?

Tick "Remove duplicates" (in Extract mode) or "Remove duplicate rows" (in Split mode). The stats line shows how many were removed.

Can I handle very large text?

Yes — the tool processes text in your browser. Extremely large inputs (tens of MB) may take a moment; use the progress feedback at the top of the page.

Is my text uploaded when I extract?

No. The text extractor processes everything locally — even sensitive logs and customer data never leave your device.

References

  1. Regular expressions — MDN Web Docs: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Guide/Regular_expressions
  2. RFC 4180 (CSV) — IETF: https://datatracker.ietf.org/doc/html/rfc4180

Related Reading