Server logs, export files, group chats and customer notes are full of useful data — but it's buried in prose. Manually copy-pasting emails or phone numbers out of a 10,000-line file is slow and error-prone. This guide shows how to extract emails, phone numbers, URLs, IP addresses and IDs with preset or custom regular expressions, split text into columns, and deduplicate the results — all in your browser.
When You Need a Text Extractor
| Scenario | What you pull out |
|---|---|
| Support tickets & chats | Email addresses, phone numbers, order IDs |
| Server / app logs | IP addresses, URLs, dates, error codes |
| CSV & spreadsheet exports | Columns split by delimiter, deduplicated rows |
| Market research | Currencies, hashtags, mentions, hex colors |
Extract with Preset Patterns
Most extraction tasks are covered by a handful of well-known shapes. The text extractor ships with presets for the most common ones:
- Email — the classic
name@domain.tldshape. - URL —
http(s)://…links. - IPv4 / IPv6 — dotted-quad and colon addresses.
- Phone — a lenient international pattern (with/without country code, dashes, parentheses).
- Number — integers and decimals, skipping IP segments.
- ID — hyphenated codes like
INV-2026-0042. - Date, Hex color, Hashtag, @mention, Currency — niche but common.
Presets are intentionally lenient: they find candidates fast, and you refine precision with a custom regex when needed.
Custom Regex for Precision
Presets cover 90% of cases; the last 10% needs your own rule. Type a pattern like \b[A-Z]{2}\d{4}\b to match exactly 2 uppercase letters followed by 4 digits, and toggle the g (global), i (case-insensitive) and m (multiline) flags. The tool validates your pattern and reports a friendly error if it's invalid.
Tip: wrap boundaries (
\bor lookarounds) in custom patterns. Without them, a pattern like\d{11}will also match the middle 11 digits of a longer number.
Deduplicate Results
Logs and exports repeat themselves constantly. Tick "Remove duplicates" and the output collapses to unique values, with the stats line showing exactly how many duplicates were dropped: Found: 342 → 58 unique · 284 duplicate(s) removed.
Split Text into Columns
When your data is already structured but separated by delimiters (CSV from a spreadsheet, TSV from a database export), the Split mode is faster than regex:
- Pick a delimiter — comma, tab, space, semicolon, pipe, newline, or a custom one.
- Review the table — rows render as a live table for a quick sanity check.
- Advanced options — remove duplicate rows, sort by any column (numeric or text, ascending/descending), or transpose rows ↔ columns.
- Copy as CSV — quoted and escaped correctly, ready for Excel or Google Sheets.
Frequently Asked Questions
Why doesn't my preset match everything I expected?
Presets are lenient but not perfect — phone formats vary wildly by country, and IDs are often project-specific. If a preset misses cases, copy it into the custom regex field and adjust, or write your own pattern.
How do I extract only unique values?
Tick "Remove duplicates" (in Extract mode) or "Remove duplicate rows" (in Split mode). The stats line shows how many were removed.
Can I handle very large text?
Yes — the tool processes text in your browser. Extremely large inputs (tens of MB) may take a moment; use the progress feedback at the top of the page.
Is my text uploaded when I extract?
No. The text extractor processes everything locally — even sensitive logs and customer data never leave your device.
References
- Regular expressions — MDN Web Docs: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Guide/Regular_expressions
- RFC 4180 (CSV) — IETF: https://datatracker.ietf.org/doc/html/rfc4180