Overview
Design, develop, and maintain automated solutions that extract data from websites, documents, and other unstructured sources. Ensure extracted data is accurate, consistent, and reliable for business needs.
What you'll do
- Design, develop, and implement robust web scraping scripts and data extraction tools.
- Extract data from online and offline sources, including websites, PDFs, and spreadsheets.
- Automate repetitive data collection tasks efficiently.
- Clean, transform, and validate extracted data to ensure high quality and consistency.
- Monitor and maintain existing scraping solutions.
- Troubleshoot issues including website changes and anti-scraping measures.
- Optimize scraping performance and reliability.
- Work closely with data analysts, developers, and business stakeholders to understand data requirements.
- Document scraping logic, data models, and processes for clarity and maintainability.
What you'll need
- Bachelor's degree in Computer Engineering, Computer Science, or Information Technology, or a Master's degree in Software Engineering.
- X+ years of experience in web scraping, data extraction, or a related field.
- Strong proficiency in a programming language commonly used for scraping, with Python highly preferred or JavaScript.
- Familiarity with relevant libraries and frameworks, including BeautifulSoup, Scrapy, Selenium, Puppeteer, or Playwright.
- Solid understanding of web technologies, including HTML, CSS, JavaScript, and HTTP/HTTPS protocols.
- Experience parsing data formats including JSON and XML.
- Excellent analytical and problem-solving skills to overcome data extraction challenges, including CAPTCHAs and dynamic content.
- Basic knowledge of data cleaning and transformation techniques.
- Proficient English language proficiency at C2 level.
Nice to have
- Experience with cloud platforms including AWS, Azure, or GCP for deploying and scaling scraping solutions.
- Familiarity with database technologies, including SQL and NoSQL, for storing and managing extracted data.
- Knowledge of regular expressions (regex) for advanced text pattern matching.
- Understanding of ethical scraping practices and data privacy regulations.
Details
- Location: Gurgaon, Haryana, India.
- Work mode: Hybrid.
- Work shift: Day Job (India).
- Employment type: Regular.
Read the full description and apply on the company’s own careers page.