Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Big Data Engineer
Location
United States
Posted
9 days ago
Salary
$100K - $150K / year
Seniority
Mid Level
Job Description
Big Data Engineer
Bright Vision Technologies
Role Description We are seeking an experienced Big Data Engineer to design, build, and operate large-scale data processing pipelines and analytics platforms on Hadoop and related big-data ecosystems. In this role you will be responsible for ingesting, transforming, and analyzing massive volumes of structured and unstructured data to support enterprise analytics, machine learning, and reporting workloads. The ideal candidate will combine deep technical expertise across the Hadoop ecosystem with strong software engineering fundamentals and a clear understanding of how to deliver reliable, performant, and cost-effective data platforms in production environments. Key Responsibilities - Design, develop, and operate end-to-end big-data pipelines on Hadoop, ingesting data from a diverse mix of relational, file-based, streaming, and API-driven sources. - Build robust ETL/ELT workflows using Apache Spark, Hive, Pig, and Sqoop, with strong attention to data quality, idempotency, error handling, and recoverability. - Develop high-throughput streaming data pipelines using Kafka, Spark Streaming, or Flink, and integrate them with downstream analytical and operational systems. - Optimize Spark and MapReduce jobs through careful tuning of partitioning, memory, serialization, and skew handling to meet demanding SLAs at minimal cost. - Design and maintain data models and storage layouts on HDFS, Hive, HBase, and modern lakehouse formats (Parquet, ORC, Delta, Iceberg, Hudi) to balance flexibility and performance. - Implement data governance, lineage, and quality controls in collaboration with data governance and security teams. - Build robust monitoring, alerting, and logging strategies for big-data pipelines, including job-level SLAs and proactive failure detection. - Partner with data scientists and analysts to deliver curated, reliable, and well-documented datasets that accelerate their work. - Automate pipeline orchestration using Airflow, Oozie, or similar workflow engines, with clean dependency management and clear ownership boundaries. - Continuously evaluate and adopt new technologies in the big-data and cloud ecosystem (EMR, Databricks, Snowflake, BigQuery) where they offer meaningful improvements. - Lead performance reviews and architecture audits of existing pipelines, proposing concrete refactoring and optimization initiatives. - Document data architectures, schemas, pipeline behaviors, and operational runbooks in a way that makes the platform supportable as the team scales. - Mentor junior engineers and contribute to the team’s engineering standards and best practices. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of professional experience designing and operating big-data pipelines on Hadoop. - Strong hands-on expertise with Apache Spark (Scala, Python, or Java) in production environments. - Solid experience with Hive, HDFS, Sqoop, HBase, and the broader Hadoop ecosystem. - Hands-on experience with streaming data platforms such as Kafka, Spark Streaming, or Flink. - Strong SQL skills and experience working with both relational and NoSQL data stores. - Experience with workflow orchestration tools such as Airflow or Oozie. - Solid understanding of distributed systems concepts, including partitioning, replication, and fault tolerance. - Strong scripting skills in Python or Shell. - Excellent troubleshooting, debugging, and documentation skills. Preferred Qualifications - Experience operating Hadoop on cloud platforms such as AWS EMR, Azure HDInsight, or Databricks. - Familiarity with modern lakehouse formats (Delta, Iceberg, Hudi). - Exposure to data governance tooling such as Apache Atlas or Collibra. - Experience with Kubernetes-based data platforms (Spark-on-K8s, Trino). - Hands-on experience with CI/CD and infrastructure-as-code in data engineering workflows. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
Data Integration & Middleware Engineer
ZensarAt Zensar, we’re “experience-led everything”. We are committed to conceptualizing, designing, engineering, marketing, and managing digital solutions and experiences for over 130 leading enterprises. We are a company driven by a bold purpose: Together, we shape experiences for better futures. Whether for our clients, our people, or the world around us, this belief powers everything we do. At the heart of our culture is ONE with Client - a set of four core values that reflect who we are and how we work: One Zensar, Nurturing, Empowering, and Client Focus. Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Role Description We are seeking a skilled Data Integration & Middleware Engineer to implement and manage scalable data pipelines and middleware solutions. The role will focus on orchestrating data workflows using Azure technologies such as Azure Data Factory (ADF), Azure Blob Storage, and Azure SQL Database, ensuring efficient ETL/ELT processing and high-quality data delivery. Qualifications - Strong experience with Azure Data Factory (ADF) for orchestration - Hands-on experience with ETL/ELT development - Experience with Azure Blob Storage / Data Lake - Proficiency in Azure SQL Database / SQL Server - Understanding of data warehousing concepts and data modeling - Knowledge of DevOps, CI/CD pipelines for data workflows - Experience with Azure Databricks or Synapse Analytics (Nice to Have) - Knowledge of data pipeline optimization and monitoring (Nice to Have) - Familiarity with API integrations and data connectors (Nice to Have) - Exposure to data governance and security practices (Nice to Have) - Azure certifications (e.g., AZ-900, DP-203) (Nice to Have) - Experience with real-time data processing frameworks (Event Hub, Service Bus) (Nice to Have) Requirements - Manage end-to-end data integration pipelines using Azure Data Factory (ADF) - Orchestrate and monitor data workflows, ensuring reliability, performance, and error handling - Implement/Manage ETL/ELT processes for data ingestion, transformation, and loading - Manage and optimize data storage solutions using Azure Blob Storage - Maintain structured data models in Azure SQL Database / Azure DB - Ensure data quality, validation, and governance across pipelines - Support integration with on-premise and cloud-based systems - Perform performance tuning and cost optimization across the data platform Benefits - Strong analytical and problem-solving abilities - Clear communication and stakeholder management skills - Ability to work in a collaborative, cross-functional environment - Attention to detail and commitment to data quality
• Lead the design and implementation of data ingestion architectures in Snowflake, ensuring scalability and reliability. • Own the development of end-to-end ELT pipelines integrating multiple data sources. • Design and enforce data quality frameworks, including validation rules, testing strategies, and monitoring. • Apply and drive Medallion Architecture (Bronze, Silver, Gold) best practices across data pipelines. • Build and optimize data pipelines and architectures on AWS (S3, Glue, Lambda, EMR, etc.). • Develop and enhance data validation and ingestion frameworks, including UI components (Streamlit-based) when needed. • Act as a technical leader, guiding best practices in data engineering, governance, and pipeline reliability. • Proactively identify risks, bottlenecks, and data quality issues, and implement mitigation strategies before they impact delivery. • Contribute to data governance initiatives, including data definitions, lineage, and stewardship practices. • Drive documentation and continuous improvement of data platform processes.
Data Scientist – Signal Processing Engineer, Acoustics
Cutsforth Inc.Truly innovative, quality products for the Power Generation Industry designed to solve problems like never before.
• Applies data science and machine learning to the analysis of electrical, vibration, and acoustic signals, transforming raw time-series sensor data into actionable diagnostics and predictive insights for rotating industrial equipment. • Partners with engineering and domain experts to design and deploy production-grade signal processing and ML solutions for predictive maintenance across industrial applications. • Operates effectively in ambiguous problem spaces where signal quality, environmental noise, and domain constraints require both technical rigor and adaptive thinking. • Design and develop signal processing pipelines and machine learning models that operate on electrical (current/voltage), vibration, and acoustic time-series sensor data, including symmetrical component analysis, matched filtering, wavelet decomposition, and time-frequency analysis techniques. • Evaluate algorithm performance using both objective metrics and subjective measures, including integration with speech recognition engines where applicable. • Perform exploratory data analysis, feature engineering, and signal feature extraction on raw electrical, vibration, and acoustic data to surface fault patterns and anomalies. • Analyze and interpret signals from electrical asset monitoring systems (motors, generators, pumps) utilizing electrical signature analysis, vibration analysis, and signal processing expertise to support fault isolation and anomaly detection. • Use cross-sensor asset monitoring data (temperature, speed, load) to characterize and validate signal-derived diagnostics. • Apply data-driven signal processing methods to characterize and isolate faults at the subsystem, component, and machine level, identifying root causes from spectral, electrical, and vibration sensor data in rotating industrial equipment. • Contribute to end-to-end ML workflows including data ingestion, model training, inference, and monitoring for drift and degradation in live environments. • Collaborate with engineering, product, and domain SMEs to translate operational challenges into well-scoped data science solutions. • Communicate findings, model performance, and business value clearly through visualizations, written documentation, and presentations to technical and non-technical stakeholders. • Explore and evaluate emerging signal processing and AI techniques, recommending production incorporation where appropriate.
• Maintain and evolve corporate data environments • Develop, maintain and optimize ETL (Extract, Transform, Load) processes • Perform database maintenance and tuning to ensure performance, availability and data integrity • Administer and support SAS environments, including tasks related to SAS DBA and SAS DI Studio (SIS) • Support legacy applications and processes, proposing continuous improvements • Monitor, analyze and address failures in data load and integration processes • Handle a significant volume of tickets, meeting established SLAs and ensuring operational continuity • Assist in identifying and resolving incidents related to the data environment • Produce and maintain technical documentation for processes and procedures • Participate in continuous improvement, automation and data architecture modernization initiatives • Support projects related to the implementation and evolution of the Data Lake


