Job Closed

This listing is no longer active.

Data Engineer (Threat Intelligence)

Location

India

Posted

80 days ago

Salary

$1.3K - $2K / month

Seniority

Mid Level

Job Description

Data Engineer (Threat Intelligence)

HEROIC Cybersecurity

Role Description HEROIC Cybersecurity (HEROIC.com) is seeking a senior-level Threat Intelligence Data Engineer – Automated Collection & Dark Web Intelligence to design, build, and operate fully automated intelligence collection systems that power our AI-driven cybersecurity and breach intelligence platforms. This role owns the end-to-end discovery, acquisition, and ingestion pipeline for continuously discovering, crawling, extracting, indexing, and normalizing millions of new artifacts daily—including documents, chats, forums, leaked datasets, repositories, threat actor communications, hacker marketplaces, unsecured infrastructure, and decentralized networks across the surface web, deep web, dark web, and anonymized networks. Our Threat Research Team’s mission is aggressive: achieve near-total coverage of global breach and leak data with 99%+ automation. Your work directly enables HEROIC’s ability to identify exposures before they are weaponized. What You Will Do - Automated Intelligence Collection & Discovery - Architect and operate large-scale, distributed crawling and discovery systems across: - Surface web, deep web, and dark web - Hacker forums, underground marketplaces, and breach communities - Chat platforms (Telegram, Discord, IRC, WhatsApp, etc.) - Paste sites, code repositories, and social platforms used for breach disclosure - Continuously discover, archive, and download newly released datasets, logs, credentials, and artifacts the moment they appear - Dark Web, Anonymized & Decentralized Networks - Build automated collectors and archivers for anonymized and decentralized networks including: - Tor (.onion), I2P, ZeroNet, Freenet, IPFS, GNUnet, Lokinet, Yggdrasil, and similar systems - Design resilient workflows for unreliable, adversarial, or ephemeral data sources - Normalize and index data from non-traditional network protocols and formats - Infrastructure & Exposure Discovery - Develop automated scanning systems to identify: - Unsecured databases (Elasticsearch, MySQL, PostgreSQL, MongoDB, etc.) - Exposed cloud storage (S3, Azure, GCP, DigitalOcean Spaces) - Open FTP servers, backups, and misconfigured archives - Monitor and ingest data from file hosting and distribution platforms commonly used for breach dumps - Pipeline Engineering & Operations - Build ETL pipelines to clean, normalize, enrich, and index structured and unstructured data - Implement advanced anti-bot evasion strategies (proxy rotation, fingerprinting, CAPTCHA mitigation, session management) - Integrate collected intelligence into centralized databases and search systems - Design APIs and internal tooling to support downstream analysis and AI/ML workflows - Automate deployment, scaling, and monitoring using Docker, Kubernetes, and cloud infrastructure - Continuously optimize performance, reliability, and cost efficiency of crawler clusters Qualifications - Minimum 4 years of hands-on experience in data engineering, intelligence collection, crawling, or distributed data pipelines - Strong Python expertise and experience with frameworks such as Scrapy, Playwright, Selenium, or custom async systems - Proven experience operating high-volume, automated data collection systems in production - Deep understanding of web protocols, HTTP, DOM parsing, and adversarial scraping environments - Experience with asynchronous, concurrent, and distributed architectures - Familiarity with SQL and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, Cassandra) - Strong Linux/Unix, shell scripting, and Git-based workflows - Experience deploying and operating systems using Docker, Kubernetes, AWS, or GCP - Excellent analytical, debugging, and problem-solving skills - Strong written and verbal communication skills Preferred / High-Value Experience - Direct experience with dark web intelligence, breach data, OSINT, or threat research - Familiarity with Tor, I2P, underground forums, stealer logs, or credential ecosystems - Experience processing large breach datasets or stealer logs - Background working in adversarial data environments - Exposure to AI/ML-driven intelligence platforms Benefits - Position Type: Full-time - Location: Remote in India. Work from wherever you please! Your home, the beach, our offices, etc. - Compensation: USD 1300-2000 monthly (depending on experience) - Professional Growth: Amazing upward mobility in a rapidly expanding company. - Innovative Culture: Be part of a team that leverages AI and cutting-edge technologies.

Related Categories

Related Job Pages

More Data Engineer Jobs

Full TimeRemoteTeam 10,001+Since 1856H1B Sponsor

• Management of Cogito suite of reporting tools • Cogito Security management • Maintenance of Cogito Related System Wide Settings • Coordinates Analytics Processes/build across all Epic Applications • Oversite of Cogito Tools build and provisioning standards • Coordinating, auditing, troubleshooting technical system functionality – including reporting queues • Represent Cogito in Epic governance groups • Coordination & Management of Cogito related projects – including Epic Releases

Alaska + 6 moreAll locations: Alaska | California | Montana | New Mexico | Oregon | Texas | Washington
$4.8K - $7.6K / year
Job Closed
Full TimeRemoteTeam 501-1,000Since 2007H1B No Sponsor

• Architect and develop large-scale, mission-critical BI and data platform solutions serving millions of users across the globe, leveraging AWS native technologies including Athena, Redshift, Glue, QuickSight, and S3. • Lead the design and implementation of robust data pipelines, data lakes, and data warehouses using modern architectures (Iceberg, Parquet, columnar formats) to support real-time and batch analytics at scale. • Drive technical strategy and architectural decisions for the BI platform, including data modeling, query optimization, performance tuning, and cost optimization across AWS services. • Build and maintain sophisticated back-end services, ETL/ELT workflows, and front-end analytics applications using Python, SQL, React, and modern web technologies. • Design and implement efficient data storage solutions across relational databases (Redshift, PostgreSQL) and non-relational databases (DynamoDB, S3), ensuring optimal performance and cost-efficiency. • Develop and maintain REST APIs and event-driven architectures to enable seamless integration between data services, analytics tools, and customer-facing applications. • Serve as the technical lead and mentor for engineering teams, conducting architecture reviews, code reviews, and providing guidance on complex technical challenges. • Collaborate with cross-functional teams including data engineers, analytics engineers, product managers, and DevOps to deliver innovative BI solutions that drive business value. • Champion engineering excellence by establishing best practices, design patterns, and coding standards for data-intensive applications. • Lead Agile ceremonies, drive sprint planning, and ensure timely delivery of high-quality software solutions while maintaining technical debt at manageable levels. • Evaluate and integrate emerging AWS services and open-source technologies to continuously improve platform capabilities and developer productivity. • Troubleshoot and resolve complex performance issues in distributed data systems, optimizing query performance, data processing workflows, and infrastructure costs. • Participate in strategic planning and roadmap development, translating business requirements into scalable technical solutions. • Contribute to the team on-call rotation, providing expert-level support for production environments and mentoring team members on incident response.

Canada
$120K / year

Senior Data Engineer

Capital Bank Maryland

Capital Bank, established in 1999, is a publicly traded company with over $2.43 billion in assets. The bank operates five branches in the greater Washington, DC

Data Engineer80 days ago

• Serve as the expert for ETL and DB solutions, collaborating with business stakeholders and IT teams to define requirements, gather data, and implement optimized data solutions. • Design, implement, and maintain data systems in Snowflake to ensure data scalability and accessibility. • Implement and manage data lakes and data warehouses, creating pipelines and data models to enable efficient analytics and reporting. • Establish and document strategies for managing data transfer processes, including secure file transfers (SFTP), batch data processing, and real-time streaming. • Build and optimize ETL pipelines for data extraction, transformation, and loading into operational databases or analytical platforms. • Integrate and support data visualization tools such as Power BI, Sisense, Google Looker, Tableau, or similar platforms to enable actionable insights for business stakeholders. • Develop and maintain optimized data models for dashboards and reporting, ensuring compatibility with visualization tools. • Plan, coordinate, and implement database migrations, upgrades, and patches with minimal downtime. • Define and enforce database governance policies, including data integrity, security, and compliance with regulatory requirements. • Analyze and resolve database performance issues by optimizing queries, indexes, and schema designs. • Partner with vendors to evaluate, select, and implement database tools, services, and technologies; stay informed about product roadmaps and industry trends. • Develop disaster recovery and high-availability solutions, including replication, clustering, and failover.

District Of Columbia + 7 moreAll locations: District Of Columbia | Florida | Illinois | North Carolina | Maryland | Pennsylvania | South Carolina | Virginia
$115K - $130K / year
Job Closed
Full TimeRemoteTeam 1,001-5,000H1B No Sponsor

• Design and implement data pipelines (ETL/ELT) using modern tools (e.g., Apache Airflow, DBT, Dataflow); • Integrate data from transactional systems, APIs, and relational and non-relational databases; • Create and maintain optimized data structures in analytical environments (data lakes and data warehouses); • Ensure data governance, data quality, and data cataloging; • Automate routines for data extraction, transformation, and loading; • Support data scientists, analysts, and product squads with reliable, well-modeled data; • Participate in modernization and data migration projects to the cloud; • Monitor and resolve failures in pipelines and other critical data processes.

Brazil
Job Closed