Simple Solutions for Complex Problems
Data Engineer
Location
United States
Posted
5 days ago
Salary
$110K - $140K / year
Seniority
Senior
Job Description
Data Engineer
Capital Technology Group, LLC
• Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), Databricks, dbt, SQL (PostgreSQL), and AWS Glue. • Develop and optimize AWS-native data platforms leveraging AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch. • Build high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro. • Design and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies. • Integrate enterprise and external data sources across relational and NoSQL platforms including PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL databases. • Build AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vector and OpenSearch vector indexes. • Develop cloud infrastructure using CloudFormation (Infrastructure as Code), GitHub, Harness, and enterprise CI/CD pipelines while leveraging SNS, SQS, and EventBridge for event-driven architectures. • Improve the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, automation, and continuous optimization. • Support mission-critical analytics and reporting solutions within large-scale AWS-based federal data environments, implementing solutions that comply with FedRAMP and NIST 800-53 security controls. • Lead modernization initiatives migrating legacy platforms including IBM DataStage, Hadoop, RunDeck, and shell-based workflows to cloud-native AWS services. • Mentor junior engineers through technical guidance, architecture discussions, and code reviews while promoting engineering best practices. • Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts effectively to technical and non-technical stakeholders.
Job Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field
- 4+ years of professional experience in data engineering or related domains
- Strong hands-on experience with: Databricks, Apache Spark (PySpark), Python, SQL (PostgreSQL), and dbt for large-scale data engineering, ETL/ELT development, data transformation, and data modeling.
- Designing, building, and maintaining AWS-native data platforms using AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, AWS Step Functions, Amazon S3, Amazon Redshift, Amazon RDS, AWS DMS, and Amazon CloudWatch.
- Developing scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments.
- Working with modern data lake technologies including Apache Iceberg and data formats such as Parquet, ORC, and Avro.
- Designing and optimizing solutions using relational and NoSQL databases including PostgreSQL, Redshift, Oracle, GraphDB, and other NoSQL platforms.
- Building reliable, high-performance data platforms through performance tuning, system optimization, and enterprise-scale ETL/ELT architectures.
- Strong analytical and problem-solving skills
- Experience working in agile, iterative software development environments
- Excellent written and verbal communication skills, with the ability to explain complex topics to diverse audiences.
Benefits
- Remote Work (Hybrid roles will be specified in the job post)
- Competitive Compensation Package
- Medical, Dental, and Vision
- Life Insurance, Short/Long Term Disability
- Employee Assistance Program
- 401(k) with 4% matching
- Liberal PTO vacation policy
- Generous Annual Continuing Education
- Annual Wellness Budget
- Bonus Incentive Programs (Employee referrals and performance-based rewards)
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
Principal Data Engineer
JobGetThe go-to marketplace for the Everyday Worker. We help employers meet job seekers where they are.
• Lead the technical direction for how data flows, scales, and powers decisions at JobGet • Own the data architecture behind the platform • Drive architectural decisions • Champion modern data stack adoption • Ensure platform reliability • Build production-grade pipelines • Establish data modeling standards • Solve complex data integration and performance challenges • Architect and evolve streaming data infrastructure • Enable machine learning at scale • Establish data governance standards • Implement data validation frameworks • Raise the technical bar through mentoring
• Collaborates with cross-functional teams to understand data and analytical requirements and translates them into effective Microsoft Fabric–based data solutions. • Mentors and provides technical leadership to Data Engineers through code reviews, design reviews, knowledge-sharing sessions, and engineering best practices. • Establishes and maintains data engineering standards, reusable frameworks, development patterns, and operational best practices to improve consistency and scalability across the data platform. • Leads root cause analysis and resolution efforts for critical production issues impacting enterprise data pipelines, reporting solutions, and platform operations. • Reviews and provides guidance on solution designs, architecture decisions, and implementation approaches to ensure alignment with enterprise standards and best practices. • Designs, develops, and maintains end-to-end data ingestion and transformation processes using Microsoft Fabric OneLake, Lakehouses, Data Pipelines, Dataflows Gen2, and Notebooks. • Formats, cleanses, and stores data in a structured manner to ensure data quality and accessibility for reporting and analysis, following Medallion Architecture. • Works closely with reporting and analytics teams to ensure datasets are accurate, well-modeled, and optimized for Power BI and other analytical workloads. • Creates comprehensive documentation for data pipelines, processes, and solutions. • Participates in data governance activities to ensure data integrity, security, and compliance. • Utilizes source control to manage and track changes to data pipelines and codebase. • Provides level II/III support to troubleshoot and resolve data-related issues and inquiries. • Provides ancillary support for Machine Learning (ML) and Artificial Intelligence (AI) processes. • Evaluates emerging technologies, industry trends, and architectural patterns related to data engineering, analytics, artificial intelligence, and cloud platforms and provides recommendations for adoption. • Manages and optimizes Microsoft Fabric platform performance by monitoring capacity utilization, query performance, Lakehouse and Warehouse workloads, Delta table maintenance (OPTIMIZE/VACUUM), data partitioning, storage consumption, and pipeline execution to ensure scalable, reliable, and cost-effective data operations. • Develops and maintains monitoring, alerting, and observability solutions for Microsoft Fabric pipelines, notebooks, Lakehouses, Warehouses, semantic models, and related data platform components to ensure operational reliability and rapid issue resolution. • Performs other duties as assigned.
• Collaborate with cross-functional teams to understand data and analytical requirements and translate them into effective Microsoft Fabric–based data solutions. • Design, develop and maintain end-to-end data ingestion and transformation processes using Microsoft Fabric Data Pipelines, Dataflows Gen2, and Notebooks. • Format, cleanse, and store data in a structured manner to ensure data quality and accessibility for reporting and analysis, following Medallion Architecture • Work closely with reporting and analytics teams to ensure datasets are accurate, well-modeled, and optimized for Power BI and other analytical workloads. • Create comprehensive developer documentation for data pipelines, processes, and solutions. • Participate in data governance activities to ensure data integrity, security, and compliance. • Utilize source control to manage and track changes to data pipelines and codebase. • Provide level II/III support to troubleshoot and resolve data-related issues and inquiries. • Provide ancillary support for Machine learning (ML) and Artificial Intelligence (AI) processes • Stay updated with industry trends and best practices related to data warehousing, ETL, and data engineering.
Data & Analytics Engineer
Eltropy Inc.Eltropy is on a mission to disrupt the way people access financial services. Eltropy enables financial institutions to digitally engage in a secure and compliant way. Using our world-class digital communications platform, community financial institutions can improve operations, engagement, and productivity. CFIs (Community Banks and Credit Unions) use Eltropy to communicate with consumers via Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology — all integrated in a single platform bolstered by AI, skill-based routing, and other contact center capabilities. Customers are our North Star No Fear - Tell the truth Team of Owners Eltropy is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Role Description Eltropy is a digital conversations platform for credit unions and community financial institutions in the US. The Data Engineering & Analytics team builds the AWS data pipelines and customer-facing dashboards that power analytics across the platform. We are looking for a Data & Analytics Engineer with 3-4 years of experience who can own dashboard delivery end to end along with the pipelines behind it. The ideal candidate learns fast, builds product context quickly, listens well, and collaborates effectively across product, engineering, DevOps, and customer-facing teams. Key Responsibilities - Own dashboard changes end to end in QuickSight and ThoughtSpot - new metrics and filters, SPICE refresh management, internal-to-production promotion, and post-release validation. - Build and maintain batch and streaming ETL pipelines on AWS using Glue (PySpark), S3, Redshift, and Airflow (MWAA) DAGs. - Write and optimize Redshift SQL; debug query performance, connection contention, and data mismatches across sources. - Support near-real-time ingestion (Kafka/MSK CDC → Glue Streaming → S3 → Redshift). - Investigate customer-reported analytics discrepancies (Jira/support tickets), root-cause them in the data, and communicate findings clearly to support, product, and engineering. - Set up and respond to pipeline monitoring - CloudWatch metrics and alarms, monitoring DAGs, refresh health - and participate in incident triage and RCA. - Develop deep product knowledge: understand what each metric means to our credit union customers and translate product changes into data model and dashboard updates. - Ensure data quality, validation, and consistency across systems. Qualifications - 3-4 years of experience in data engineering and/or analytics engineering. - Strong SQL on Redshift (or a similar MPP warehouse) and solid Python/PySpark. - Hands-on experience with the AWS data stack: S3, Glue, Redshift, CloudWatch. - Workflow orchestration with Apache Airflow — authoring, debugging, and deploying DAGs. - Data modeling and warehousing fundamentals. - BI dashboarding experience with QuickSight, ThoughtSpot, Tableau, or Power BI. - Quick learner — able to grasp an unfamiliar product and data model fast and work independently. - Strong listening and collaboration skills; comfortable coordinating across product, engineering, DevOps, and customer-facing teams. Requirements - Streaming/CDC experience: Kafka (MSK), Debezium, Spark Structured Streaming. - AWS infrastructure awareness: IAM roles, security groups, Secrets Manager, Kinesis/Firehose. - Fintech or B2B SaaS analytics exposure. Security Responsibilities - Adhere to Eltropy's policies on security, confidentiality, availability, and privacy; handle customer and financial-institution data responsibly, protect access credentials, and report security events promptly through Eltropy's channels.


