முக்கிய விஷயங்கள்
- Regex is just a mini pattern-matching language - not magic, just rules you can learn in an afternoon.
- Character classes like \d, \w and [A-Z] decide what counts as a valid character in your input.
- Anchors (^ $) and quantifiers (+, *, {n,m}) control position and how many times something repeats.
- Capture groups () let you pull out just the PAN, PIN code or OTP you need, not the whole matched string.
- Indian mobile numbers, PAN and PIN codes each need a slightly different pattern - don't copy-paste blindly from Stack Overflow.
What just happened?
Picture this. You built a signup form for your startup's app last week. Someone typed rahul@@gmail..com into the email field, and your backend happily saved it. Three weeks later, your transactional emails are bouncing and nobody knows why. This happens in almost every Indian product team, from a two-person Chennai startup to a floor at a Bangalore unicorn. The fix usually isn't more code - it's one good regular expression sitting at the validation layer, catching garbage before it touches your database.
Regex has a reputation for looking scary. Honestly, that line of backslashes and brackets does look like someone fell asleep on the keyboard. But once you know five or six building blocks, you can read almost any regex you'll meet in a real Indian codebase - phone number checks, PAN validators, PIN code lookups, log file parsers.
How does this actually work?
Let's start with the basics - characters and classes. A regex engine reads your pattern character by character and tries to match it against the input text. Some characters match themselves literally. Others, called classes, match a whole category of characters.
\d matches any digit, \w matches any letter, digit or underscore, and \s matches a space or tab. You can also build your own class with square brackets, like [A-Z] for any uppercase letter.
// Character classes in action
const input = "My OTP is 482913";
const digitsOnly = input.match(/\d+/);
console.log(digitsOnly[0]); // 482913Next comes quantifiers and anchors - the part that decides how many times a character can repeat, and where in the string it must sit. + means one or more, * means zero or more, and {n,m} means between n and m times. Anchors ^ and $ pin the match to the start and end of the string, which matters a lot once your input has extra spaces or hidden characters.
// Quantifiers + anchors - no stray text allowed
const pattern = /^[6-9]\d{9}$/;
console.log(pattern.test("9876543210")); // true
console.log(pattern.test("98765432100")); // false - 11 digits
console.log(pattern.test("abc9876543210")); // false - anchors block itNow, groups and capture - this is where regex stops being just a yes/no check and starts being useful for extraction. Wrap part of your pattern in round brackets () and that chunk becomes a capture group you can pull out separately.
// Capture groups - pulling out just the useful bits
const panRegex = /^([A-Z]{5})(\d{4})([A-Z])$/;
const result = "ABCDE1234F".match(panRegex);
console.log(result[1]); // ABCDE - name-based letters
console.log(result[2]); // 1234 - serial number
console.log(result[3]); // F - check characterThat's really the whole toolkit - characters, quantifiers, anchors, groups. Everything else in regex is just combinations of these four ideas.
What changes for people in India?
Indian data has its own quirks, and generic tutorials rarely cover them. A 10-digit Indian mobile number always starts with 6, 7, 8 or 9, which is why the character class [6-9] showed up above instead of a plain \d.
PAN numbers follow a fixed, documented structure - five letters, four digits, one letter - which is exactly why the capture group example above works cleanly without any guesswork. PIN codes are simpler on the surface but easy to get wrong: six digits, and the first digit should never be zero.
// Common Indian validators in one place
const mobileRegex = /^[6-9]\d{9}$/;
const panRegex = /^[A-Z]{5}\d{4}[A-Z]$/;
const pinCodeRegex = /^[1-9]\d{5}$/;
console.log(mobileRegex.test("9123456789")); // true
console.log(panRegex.test("AAAPL1234C")); // true
console.log(pinCodeRegex.test("600028")); // true, Chennai PINThese three patterns alone cover a huge chunk of what Flipkart, Swiggy, Paytm or any UPI-linked app checks before saving your profile. The email regex most teams actually use in production is intentionally loose - something like /^[^\s@]+@[^\s@]+\.[^\s@]+$/ - because fully validating RFC-compliant emails with regex alone is a rabbit hole nobody finishes on time.
What should you do now?
Here's where most developers trip up, and it's worth naming the mistakes plainly.
Common mistakes: forgetting anchors, so /\d{10}/ happily matches inside a 15-digit string instead of rejecting it. Using greedy quantifiers like .* when you meant a narrower, specific class, which slows down matching on long text and sometimes matches way more than you intended. Testing only the happy path - nobody tries an empty string, a string with just spaces, or Unicode characters until it breaks in production. And trusting a regex you don't understand, copied from a forum post, without running your own test cases first.
In Indian companies, this comes up constantly - at TCS, Infosys or any product startup, form validation and data-cleaning scripts lean on regex daily. It's also a near-guaranteed interview question: write a regex to validate an email or extract digits from a mixed string is asked at junior and mid-level developer interviews across the country, because it tests whether you actually understand pattern logic or just memorised one line.
Practice task: Take a messy CSV column of customer support numbers with inconsistent formatting - some with +91, some with spaces, some with dashes. Write one regex that extracts just the 10-digit number from each row, ignoring the country code and formatting characters. Test it against at least five different formats before you trust it.
Where does this regex actually sit in your stack?
Most junior developers assume one regex check on the frontend is enough, but that's rarely how production apps work in India. A Flipkart or Swiggy signup form validates the phone number in the browser for instant feedback, then validates the exact same pattern again on the server before anything touches the database. Skip that second check and you're trusting a browser that a user can bypass with one click of dev tools or a direct API call from Postman. The rule senior engineers repeat at every Chennai or Bangalore standup is simple: client-side regex is for user experience, server-side regex is for data integrity, and you never skip the second one just because the first one passed.
There's also a real performance trap hiding inside badly written regex, something called catastrophic backtracking. A pattern like /(a+)+b/ looks harmless but can make the regex engine try thousands of combinations on a long string with no matching 'b', freezing your Node.js server for seconds on what looks like a tiny validation call. This has genuinely taken down production APIs, including payment and OTP flows, when someone pasted a 'smart-looking' regex from a blog without testing it against long or adversarial input. If your validator runs on every UPI transaction or every OTP request, a slow regex isn't a minor bug - it's a bottleneck your whole checkout flow feels.
What's worth watching going forward is how much regex tooling has improved without changing the core syntax you learned above. Named capture groups like (?<mobile>[6-9]\d{9}) make long patterns far more readable in code reviews, and most modern JavaScript, Python and Java runtimes support them cleanly now. Online testers like regex101 remain the fastest way to debug a pattern before it ships, letting you paste real Indian sample data - mixed PAN formats, PIN codes, messy phone numbers with +91 prefixes - and see exactly which part of your pattern is failing, instead of guessing and redeploying five times.




Comments (0)
Be the first to comment!