Threat Intelligence Data Engineer
Location
India
Posted
84 days ago
Salary
$1.3K - $2K / month
Seniority
Mid Level
Job Description
Threat Intelligence Data Engineer
HEROIC.com
Role Description HEROIC Cybersecurity (HEROIC.com) is seeking a senior-level Threat Intelligence Data Engineer - Automated Collection & Dark Web Intelligence to design, build, and operate fully automated intelligence collection systems that power our AI-driven cybersecurity and breach intelligence platforms. This role owns the end-to-end discovery, acquisition, and ingestion pipeline for continuously discovering, crawling, extracting, indexing, and normalizing millions of new artifacts daily—including documents, chats, forums, leaked datasets, repositories, threat actor communications, hacker marketplaces, unsecured infrastructure, and decentralized networks across the surface web, deep web, dark web, and anonymized networks. Our Threat Research Team’s mission is aggressive: achieve near-total coverage of global breach and leak data with 99%+ automation. Your work directly enables HEROIC’s ability to identify exposures before they are weaponized. What You Will Do - Automated Intelligence Collection & Discovery - Architect and operate large-scale, distributed crawling and discovery systems across: - Surface web, deep web, and dark web - Hacker forums, underground marketplaces, and breach communities - Chat platforms (Telegram, Discord, IRC, WhatsApp, etc.) - Paste sites, code repositories, and social platforms used for breach disclosure - Continuously discover, archive, and download newly released datasets, logs, credentials, and artifacts the moment they appear - Dark Web, Anonymized & Decentralized Networks - Build automated collectors and archivers for anonymized and decentralized networks including: - Tor (.onion), I2P, ZeroNet, Freenet, IPFS, GNUnet, Lokinet, Yggdrasil, and similar systems - Design resilient workflows for unreliable, adversarial, or ephemeral data sources - Normalize and index data from non-traditional network protocols and formats - Infrastructure & Exposure Discovery - Develop automated scanning systems to identify: - Unsecured databases (Elasticsearch, MySQL, PostgreSQL, MongoDB, etc.) - Exposed cloud storage (S3, Azure, GCP, DigitalOcean Spaces) - Open FTP servers, backups, and misconfigured archives - Monitor and ingest data from file hosting and distribution platforms commonly used for breach dumps - Pipeline Engineering & Operations - Build ETL pipelines to clean, normalize, enrich, and index structured and unstructured data - Implement advanced anti-bot evasion strategies (proxy rotation, fingerprinting, CAPTCHA mitigation, session management) - Integrate collected intelligence into centralized databases and search systems - Design APIs and internal tooling to support downstream analysis and AI/ML workflows - Automate deployment, scaling, and monitoring using Docker, Kubernetes, and cloud infrastructure - Continuously optimize performance, reliability, and cost efficiency of crawler clusters Qualifications - Minimum 4 years of hands-on experience in data engineering, intelligence collection, crawling, or distributed data pipelines - Strong Python expertise and experience with frameworks such as Scrapy, Playwright, Selenium, or custom async systems - Proven experience operating high-volume, automated data collection systems in production - Deep understanding of web protocols, HTTP, DOM parsing, and adversarial scraping environments - Experience with asynchronous, concurrent, and distributed architectures - Familiarity with SQL and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, Cassandra) - Strong Linux/Unix, shell scripting, and Git-based workflows - Experience deploying and operating systems using Docker, Kubernetes, AWS, or GCP - Excellent analytical, debugging, and problem-solving skills - Strong written and verbal communication skills Preferred / High-Value Experience - Direct experience with dark web intelligence, breach data, OSINT, or threat research - Familiarity with Tor, I2P, underground forums, stealer logs, or credential ecosystems - Experience processing large breach datasets or stealer logs - Background working in adversarial data environments - Exposure to AI/ML-driven intelligence platforms Benefits - Position Type: Full-time - Location: Remote in India. Work from wherever you please! Your home, the beach, our offices, etc. - Compensation: USD 1300-2000 monthly - Professional Growth: Amazing upward mobility in a rapidly expanding company. - Innovative Culture: Be part of a team that leverages AI and cutting-edge technologies.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Senior Data Engineer • Experience with Databricks • Experience with Python/Spark • Experience with Unity Catalog • CI/CD pipeline deployments • Experience working in Scrum
• Contact clients directly immediately after the sale to welcome them and begin follow-up for the migration process; • Act as the frontline for starting migrations, recording and routing complaints, and coordinating with clients in complex cases; • Make calls to explain the process, resolve support questions, respond to emails, and ensure clarity in communications; • Proactively follow up with clients to request pending information essential for migration (e.g., sending backups, registering professionals, validation responses); • Receive and handle requests from the Sales, Support, or Customer Success (CS) teams related to migration status; • Actively research how to request backups from competing systems to guide clients and, subsequently, perform these requests independently; • Collect received backups and maintain strict organization of our migrations drive; • Perform basic data imports or imports that already have a fully established process; • Assist with migration tests to pre-validate data before final delivery to the client; • Continuously update the migration status control spreadsheet, ensuring the entire team has visibility into the progress of each clinic.
Senior Data Engineer
iFoodNo iFood, acreditamos na força da diversidade para gerar #Inovação e atingir #Resultados incríveis, por isso, não fazemos distinção para candidatos com deficiência, gênero, orientação sexual, raça/etnia, idade, origem, constituição familiar e estética. Temos grupos compostos por foodlovers voluntários, onde falamos sobre Raça, Gênero, LGBTQI+ e PcD. Queremos ser a empresa onde pessoas escolham como lugar para se desenvolver e contribuir para a realização de sonhos, #AllTogether. Nós, FoodLovers, temos fome de inovação e resultado. Buscamos sempre fazer o nosso melhor, pensando "fora da caixa" e atuando com agilidade e responsabilidade! Temos fome de diversidade, conhecimento e compartilhamento. Trabalhamos em um ambiente de muita versatilidade. Sabe o que promove a nossa receita especial? As pessoas! Vem fazer parte disso🤝
Role Description Colaborar com membros da equipe para coletar requisitos de negócio e projetar modelos de dados. - Construir confiança em todas as interações e fornecer os dados mais completos, confiáveis e precisos para a empresa. - Projetar, desenvolver e estender código ETL e pipelines de dados para expandir o Data Lake Empresarial. - Criar e manter documentação de arquitetura e sistemas. - Manter conjuntos de dados principais ou produtos de dados atualizados no Catálogo de Dados para apoiar análises de Self-Service e Fonte Única da Verdade. - Desenvolver código que atenda aos nossos padrões internos de estilo, manutenibilidade e melhores práticas para um ambiente de dados de alta escala. - Manter e defender esses padrões através de revisão de código. - Aprovar mudanças em modelos de dados como Revisor da Equipe de Dados e ser responsável por conjuntos de dados específicos e esquemas de modelos de dados. - Compartilhar expertise em modelagem e transformação de dados para todas as equipes do iFood através de revisões de código, pareamento e treinamento para ajudar a entregar designs e consultas otimizados e escaláveis de conjuntos de dados ou produtos de dados no Databricks. Qualifications - Disposição para aprender coisas novas todos os dias. - Habilidade com SQL, Python e Spark. - Excelente comunicação: Alcançar consenso regularmente entre equipes técnicas e de negócio. - Capacidade demonstrada de comunicar de forma clara e concisa atividades complexas de negócio, requisitos técnicos e recomendações. - 2+ anos na área de Dados como analista, engenheiro ou equivalente. - 2+ anos de experiência projetando, implementando, operando e estendendo modelos de dados empresariais (dimensionais ou não). - 2+ anos trabalhando com Banco de Dados ou Data Warehouse de grande escala, preferencialmente em ambiente de nuvem. - Inglês técnico. - Desejável: Experiência com Databricks, Airflow e AWS. Requirements - Experiência prática em inteligência artificial. - Habilidade para gerar indicadores mensuráveis e identificar mudanças necessárias. - Capacidade de se adaptar rapidamente a novas tecnologias e desafios. - Interesse em promover a cultura de colaboração e troca de conhecimentos. Benefits - Buscamos uma pessoa apaixonada por tecnologia e dados, que esteja sempre em busca de novos aprendizados e que goste de desafios.
Principal Data Architect – Battery Storage
Plus PowerPlus Power develops battery energy storage systems that enable a more efficient and reliable electrical grid.
• Define and evolve data architecture standards for analytics and reporting, including data modeling, naming conventions, schema design, and documentation practices across the organization • Own the data catalog and metadata strategy, partnering with stakeholders to define, name, and organize data assets across multiple domains and source systems • Collaborate closely with Principal Data Engineering leadership and application engineering teams to align on ELT patterns, Snowflake usage, schema evolution, and analytical data modeling practices • Contribute hands‑on through SQL and Python, developing reference data models, prototypes, templates, and example implementations that demonstrate architectural intent • Support and enable data analysts by establishing consistent data usage, modeling standards, and shared definitions across a wide range of technical skill levels • Partner with application engineers on schema design to support rapid application development and reliable integration between operational and analytical data systems • Support PostgreSQL (including AWS Aurora) and Snowflake data modeling and analytical access patterns in collaboration with platform and database stakeholders • Establish and promote data governance practices covering data quality, ownership, lifecycle management, and schema change management • Drive incremental delivery of data architecture improvements, aligning short‑term progress with a clear long‑term architectural vision • Design high-ingestion pipelines (using tools like InfluxDB, Timescale, or Snowflake) capable of handling millions of data points per second from globally distributed battery sites • Ensure data can be seamlessly ingested from various industrial protocols such as Modbus, CAN bus, or DNP3, and translated into standardized cloud formats • Ensure data architectures comply with grid-specific regulations (like NERC CIP) and mandate on-site data storage for grid resilience • Help set the vision, roadmap and communicate the enterprise data strategy for the company



