‹ Back to Home

Regex Explained for Developers: Validate Phone Numbers, Emails and PAN Without Guessing

Every Indian developer writes a phone number or PAN validator at some point - and most of us copy-paste a regex we don't fully understand. Here's how characters, quantifiers, anchors and groups actually work, with Indian examples you can run today.

Keerthika 7 min read
Follow on Google
Programming Regex Explained for Developers: Validate Phone Numbers, Emails and PAN Without Guessing 7 min left Follow on Google

TamilTech AI summary

  • Four building blocks - characters, quantifiers, anchors, groups - explain almost every regex you'll meet
  • Ready-to-use regex for Indian mobile numbers, PAN and PIN codes with runnable code
  • Covers the exact mistakes and interview questions Indian dev teams actually run into

AI-assisted summary, checked by the TamilTech editorial team.

0:00
0:00
🔒 Listen is for subscribers. Subscribe

முக்கிய விஷயங்கள்

  • Regex is just a mini pattern-matching language - not magic, just rules you can learn in an afternoon.
  • Character classes like \d, \w and [A-Z] decide what counts as a valid character in your input.
  • Anchors (^ $) and quantifiers (+, *, {n,m}) control position and how many times something repeats.
  • Capture groups () let you pull out just the PAN, PIN code or OTP you need, not the whole matched string.
  • Indian mobile numbers, PAN and PIN codes each need a slightly different pattern - don't copy-paste blindly from Stack Overflow.

What just happened?

Picture this. You built a signup form for your startup's app last week. Someone typed rahul@@gmail..com into the email field, and your backend happily saved it. Three weeks later, your transactional emails are bouncing and nobody knows why. This happens in almost every Indian product team, from a two-person Chennai startup to a floor at a Bangalore unicorn. The fix usually isn't more code - it's one good regular expression sitting at the validation layer, catching garbage before it touches your database.

Regex has a reputation for looking scary. Honestly, that line of backslashes and brackets does look like someone fell asleep on the keyboard. But once you know five or six building blocks, you can read almost any regex you'll meet in a real Indian codebase - phone number checks, PAN validators, PIN code lookups, log file parsers.

How does this actually work?

Let's start with the basics - characters and classes. A regex engine reads your pattern character by character and tries to match it against the input text. Some characters match themselves literally. Others, called classes, match a whole category of characters.

\d matches any digit, \w matches any letter, digit or underscore, and \s matches a space or tab. You can also build your own class with square brackets, like [A-Z] for any uppercase letter.

// Character classes in action
const input = "My OTP is 482913";
const digitsOnly = input.match(/\d+/);
console.log(digitsOnly[0]); // 482913

Next comes quantifiers and anchors - the part that decides how many times a character can repeat, and where in the string it must sit. + means one or more, * means zero or more, and {n,m} means between n and m times. Anchors ^ and $ pin the match to the start and end of the string, which matters a lot once your input has extra spaces or hidden characters.

// Quantifiers + anchors - no stray text allowed
const pattern = /^[6-9]\d{9}$/;
console.log(pattern.test("9876543210")); // true
console.log(pattern.test("98765432100")); // false - 11 digits
console.log(pattern.test("abc9876543210")); // false - anchors block it

Now, groups and capture - this is where regex stops being just a yes/no check and starts being useful for extraction. Wrap part of your pattern in round brackets () and that chunk becomes a capture group you can pull out separately.

// Capture groups - pulling out just the useful bits
const panRegex = /^([A-Z]{5})(\d{4})([A-Z])$/;
const result = "ABCDE1234F".match(panRegex);
console.log(result[1]); // ABCDE - name-based letters
console.log(result[2]); // 1234 - serial number
console.log(result[3]); // F - check character

That's really the whole toolkit - characters, quantifiers, anchors, groups. Everything else in regex is just combinations of these four ideas.

What changes for people in India?

Indian data has its own quirks, and generic tutorials rarely cover them. A 10-digit Indian mobile number always starts with 6, 7, 8 or 9, which is why the character class [6-9] showed up above instead of a plain \d.

PAN numbers follow a fixed, documented structure - five letters, four digits, one letter - which is exactly why the capture group example above works cleanly without any guesswork. PIN codes are simpler on the surface but easy to get wrong: six digits, and the first digit should never be zero.

// Common Indian validators in one place
const mobileRegex = /^[6-9]\d{9}$/;
const panRegex = /^[A-Z]{5}\d{4}[A-Z]$/;
const pinCodeRegex = /^[1-9]\d{5}$/;

console.log(mobileRegex.test("9123456789")); // true
console.log(panRegex.test("AAAPL1234C")); // true
console.log(pinCodeRegex.test("600028")); // true, Chennai PIN

These three patterns alone cover a huge chunk of what Flipkart, Swiggy, Paytm or any UPI-linked app checks before saving your profile. The email regex most teams actually use in production is intentionally loose - something like /^[^\s@]+@[^\s@]+\.[^\s@]+$/ - because fully validating RFC-compliant emails with regex alone is a rabbit hole nobody finishes on time.

What should you do now?

Here's where most developers trip up, and it's worth naming the mistakes plainly.

Common mistakes: forgetting anchors, so /\d{10}/ happily matches inside a 15-digit string instead of rejecting it. Using greedy quantifiers like .* when you meant a narrower, specific class, which slows down matching on long text and sometimes matches way more than you intended. Testing only the happy path - nobody tries an empty string, a string with just spaces, or Unicode characters until it breaks in production. And trusting a regex you don't understand, copied from a forum post, without running your own test cases first.

In Indian companies, this comes up constantly - at TCS, Infosys or any product startup, form validation and data-cleaning scripts lean on regex daily. It's also a near-guaranteed interview question: write a regex to validate an email or extract digits from a mixed string is asked at junior and mid-level developer interviews across the country, because it tests whether you actually understand pattern logic or just memorised one line.

Practice task: Take a messy CSV column of customer support numbers with inconsistent formatting - some with +91, some with spaces, some with dashes. Write one regex that extracts just the 10-digit number from each row, ignoring the country code and formatting characters. Test it against at least five different formats before you trust it.

Where does this regex actually sit in your stack?

Most junior developers assume one regex check on the frontend is enough, but that's rarely how production apps work in India. A Flipkart or Swiggy signup form validates the phone number in the browser for instant feedback, then validates the exact same pattern again on the server before anything touches the database. Skip that second check and you're trusting a browser that a user can bypass with one click of dev tools or a direct API call from Postman. The rule senior engineers repeat at every Chennai or Bangalore standup is simple: client-side regex is for user experience, server-side regex is for data integrity, and you never skip the second one just because the first one passed.

There's also a real performance trap hiding inside badly written regex, something called catastrophic backtracking. A pattern like /(a+)+b/ looks harmless but can make the regex engine try thousands of combinations on a long string with no matching 'b', freezing your Node.js server for seconds on what looks like a tiny validation call. This has genuinely taken down production APIs, including payment and OTP flows, when someone pasted a 'smart-looking' regex from a blog without testing it against long or adversarial input. If your validator runs on every UPI transaction or every OTP request, a slow regex isn't a minor bug - it's a bottleneck your whole checkout flow feels.

What's worth watching going forward is how much regex tooling has improved without changing the core syntax you learned above. Named capture groups like (?<mobile>[6-9]\d{9}) make long patterns far more readable in code reviews, and most modern JavaScript, Python and Java runtimes support them cleanly now. Online testers like regex101 remain the fastest way to debug a pattern before it ships, letting you paste real Indian sample data - mixed PAN formats, PIN codes, messy phone numbers with +91 prefixes - and see exactly which part of your pattern is failing, instead of guessing and redeploying five times.

FAQs

Frequently asked questions

What's the actual difference between validating and extracting data with regex?

Validating just answers true or false - does the whole string match the pattern or not, using methods like test(). Extracting uses capture groups and match() to pull out specific pieces, like just the PAN serial number out of a full PAN string, even if the rest of the string also matched.

Why does my regex work on an online tester but fail in my actual code?

Usually it's a flavor mismatch or escaping issue. Online testers sometimes default to a different regex engine than your language uses, and backslashes often need double-escaping inside a string literal in languages like Java or PHP. Always test the final string your code will actually run, not just the raw pattern.

Is regex alone enough to fully validate an email address?

No, and most senior developers will tell you not to try. A loose regex that checks for an @ symbol and a dot after it catches 95% of typos. The real confirmation that an email works is sending a verification link - that's what Flipkart, Swiggy and most Indian apps actually rely on.

Which regex syntax should I learn first - JavaScript, Python or Java?

The core syntax - character classes, quantifiers, anchors, groups - is nearly identical across all three. Learn it once in whichever language you use daily, and switching to another language later is mostly about syntax for escaping and function names, not relearning the logic.

How do I test a regex before pushing it to production?

Write a small list of test cases that should pass and a list that should fail, including edge cases like empty strings, extra spaces, and slightly wrong formats. Run both lists against your pattern every time you change it, instead of eyeballing one or two examples.

Get tomorrow’s tech news on WhatsApp

One short update a day, free. Follow the TamilTech channel.

What do you think?

people reacted

Keerthika

TamilTech editorial team · 3,412 articles

Keerthika is an editor at TamilTech, the Tamil and English technology publication founded by Praveen Kumar S. She covers AI, smartphones, gadgets, EVs, startups and cybersecurity i...

More from Keerthika

Ask TamilTech on WhatsApp

Tech doubt? Ask in Tamil or English — our WhatsApp assistant answers from TamilTech articles in seconds.

Related stories

Comments (0)

| Supports **bold**, *italic*, `code`

Be the first to comment!

Next story Cursor AI Makes Massive India Push Ahead of SpaceX Acquisition: Localized Pricing & 2026 Roadmap Explained
Tamiltech

Tamiltech

Install app for faster access

Earn XP 🏆
WhatsApp
Notifications