ZipNom
Remote Jobs
1 Jobs
Role Description This is a remote position. We are looking for a Data Engineer to own and build the MCA Corporate Filings Intelligence Pipeline — the system that converts unstructured government filings (PDFs, XBRL, scanned documents, and web data) into clean, structured, queryable business intelligence. MCA does not provide an official public API. This role requires building a resilient data acquisition system using a combination of unofficial endpoints, web scraping, third-party enrichment APIs, and AI-based document extraction — and making it run reliably, every single day, without breaking. If you enjoy the challenge of "the data is out there, but it's a mess — go make it useful," this role is built for you. What You Will Build - Daily Company Ingestion Pipeline: Build and maintain the pipeline that fetches newly incorporated companies (Private Limited, LLP, OPC) from MCA every day. - Director and Contact Enrichment: Enrich every new company with director details and, where possible, director mobile numbers and emails. - Financial Filings Extraction Pipeline: Build the system that downloads AOC-4, MGT-7, CHG-1, DIR-12, and PAS-3 filings, and extracts structured financial data from them. - Data Transformation and Intelligence Layer: Normalize extracted financial data, compute financial ratios, generate a financial health score per company, and detect business signals. - Director Network Graph: Build and maintain a graph of directors-to-companies relationships. - Pipeline Orchestration and Monitoring: Schedule and monitor all jobs using AWS Step Functions / Bull queues with cron scheduling. - Data Quality and Compliance: Build validation rules, quality scoring, duplicate detection, and DPDPA-compliant data handling. What Makes This Role Interesting - You are solving a real puzzle, not following a spec. - Your output is the product. - You will work with cutting-edge AI extraction. - High ownership, fast feedback loops. A Typical Week Might Include - Investigating why the MCA unofficial API started returning 403s and adding a Selenium fallback. - Writing a new Claude extraction prompt for CHG-1 filings and validating accuracy against sample documents. - Tuning the financial health score weights after reviewing a month of computed scores. - Adding a new enrichment provider to the director mobile fallback chain. - Debugging a spike in validation errors and tracing it back to a currency-unit detection bug. - Reviewing CAPTCHA-solving costs and optimizing the caching layer. Interview Process - Initial screening call (30 minutes) — background, experience, and role fit. - Technical round (60 to 90 minutes) — data pipeline design discussion + live problem-solving. - Take-home task — a small real-world extraction or pipeline design problem. - Final round with founder — architecture discussion, culture fit, and Q&A. How to Apply Send your resume and GitHub/portfolio link to careers@verizol.ai with the subject line "Data Engineer Application — [Your Name]". Include a short note on a data pipeline or scraping project you have built. Qualifications - 2+ years of experience building data pipelines — ETL/ELT systems, scheduled jobs, or similar. - Strong Node.js or Python skills (Node.js/TypeScript required). - Solid PostgreSQL experience — schema design, indexing, writing and optimizing complex queries. - Experience with async job processing — queues, cron, retries, and failure handling. - Experience working with external APIs — authentication, rate limiting, pagination, error handling. - Strong debugging mindset — comfortable diagnosing why a pipeline silently produced bad data. - Attention to data quality — you care about validating, not just moving, data. Requirements - Language/Runtime: Node.js, TypeScript. - Database: PostgreSQL (AWS Aurora). - Queues/Orchestration: Redis, Bull, AWS Step Functions, EventBridge, Lambda. - Web Access: Axios, Selenium/Puppeteer, rotating proxies, cookie-jar session management. - CAPTCHA Solving: 2Captcha API integration. - Document Processing: pdf-parse, pdf2pic, Tesseract OCR, xml2js (XBRL parsing). - AI Extraction: Claude API (Anthropic). - Storage: AWS S3 (raw document archive). - Enrichment APIs: Sandbox.co.in, CompData, Apollo.io, GST data cross-reference. - Monitoring: CloudWatch, Sentry, WhatsApp (WATI) alerting. Benefits - Compensation: ₹8,00,000 to ₹16,00,000/year based on experience and interview performance. - ESOPs for early team members — meaningful equity in a growing company. - Direct collaboration with the founder on pipeline architecture and prioritization. - Budget allocated for third-party enrichment APIs, proxies, and CAPTCHA solving. - Flexible working hours once ramped up.