Simple Solutions for Complex Problems
Junior Data Engineer
Location
United States
Posted
1 day ago
Salary
$75K - $115K / year
Seniority
Junior
Job Description
Junior Data Engineer
Capital Technology Group, LLC
• Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Databricks, dbt, Apache Spark (PySpark), SQL (PostgreSQL), and Python. • Develop and optimize AWS-native data solutions leveraging services including AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch. • Build high-performance data ingestion, transformation, and orchestration workflows across structured and semi-structured data using Parquet, ORC, Avro, and Apache Iceberg. • Integrate data from enterprise and external sources, including relational and NoSQL databases such as PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL platforms. • Improve the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, and continuous optimization. • Develop and maintain automated deployment pipelines using Harness and collaborate on cloud infrastructure and platform improvements. • Support data engineering efforts powering mission-critical analytics, reporting, and decision-making across large-scale federal data environments. • Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and continuously improve engineering processes.
Job Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field
- 5 years of professional experience in data engineering or related domains
- Strong hands-on experience with: Databricks for large-scale data engineering, ETL/ELT development, and data transformation.
- Strong proficiency with Dbt, SQL (PostgreSQL), Python, Java, AWS (Redshift, RDS, S3, Glue, Lambda, DMS, CloudWatch), Data Pipelines, Data Modeling, System Maintenance, Data Integration, Performance Optimization.
- AWS native.
- Strong analytical and problem-solving skills
- Experience working in agile, iterative software development environments
- Ability to quickly learn and apply new technologies and domain knowledge
- Excellent written and verbal communication skills, with the ability to explain complex topics to diverse audiences.
Benefits
- Remote Work
- Competitive Compensation Package
- Medical, Dental, and Vision
- Life Insurance, Short/Long Term Disability
- Employee Assistance Program
- 401(k) with 4% matching
- Liberal PTO vacation policy
- Generous Annual Continuing Education
- Annual Wellness Budget
- Bonus Incentive Programs
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), Databricks, dbt, SQL (PostgreSQL), and AWS Glue. • Develop and optimize AWS-native data platforms leveraging AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch. • Build high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro. • Design and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies. • Integrate enterprise and external data sources across relational and NoSQL platforms including PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL databases. • Build AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vector and OpenSearch vector indexes. • Develop cloud infrastructure using CloudFormation (Infrastructure as Code), GitHub, Harness, and enterprise CI/CD pipelines while leveraging SNS, SQS, and EventBridge for event-driven architectures. • Improve the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, automation, and continuous optimization. • Support mission-critical analytics and reporting solutions within large-scale AWS-based federal data environments, implementing solutions that comply with FedRAMP and NIST 800-53 security controls. • Lead modernization initiatives migrating legacy platforms including IBM DataStage, Hadoop, RunDeck, and shell-based workflows to cloud-native AWS services. • Mentor junior engineers through technical guidance, architecture discussions, and code reviews while promoting engineering best practices. • Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts effectively to technical and non-technical stakeholders.
Principal Data Engineer
JobGetThe go-to marketplace for the Everyday Worker. We help employers meet job seekers where they are.
• Lead the technical direction for how data flows, scales, and powers decisions at JobGet • Own the data architecture behind the platform • Drive architectural decisions • Champion modern data stack adoption • Ensure platform reliability • Build production-grade pipelines • Establish data modeling standards • Solve complex data integration and performance challenges • Architect and evolve streaming data infrastructure • Enable machine learning at scale • Establish data governance standards • Implement data validation frameworks • Raise the technical bar through mentoring
• Collaborates with cross-functional teams to understand data and analytical requirements and translates them into effective Microsoft Fabric–based data solutions. • Mentors and provides technical leadership to Data Engineers through code reviews, design reviews, knowledge-sharing sessions, and engineering best practices. • Establishes and maintains data engineering standards, reusable frameworks, development patterns, and operational best practices to improve consistency and scalability across the data platform. • Leads root cause analysis and resolution efforts for critical production issues impacting enterprise data pipelines, reporting solutions, and platform operations. • Reviews and provides guidance on solution designs, architecture decisions, and implementation approaches to ensure alignment with enterprise standards and best practices. • Designs, develops, and maintains end-to-end data ingestion and transformation processes using Microsoft Fabric OneLake, Lakehouses, Data Pipelines, Dataflows Gen2, and Notebooks. • Formats, cleanses, and stores data in a structured manner to ensure data quality and accessibility for reporting and analysis, following Medallion Architecture. • Works closely with reporting and analytics teams to ensure datasets are accurate, well-modeled, and optimized for Power BI and other analytical workloads. • Creates comprehensive documentation for data pipelines, processes, and solutions. • Participates in data governance activities to ensure data integrity, security, and compliance. • Utilizes source control to manage and track changes to data pipelines and codebase. • Provides level II/III support to troubleshoot and resolve data-related issues and inquiries. • Provides ancillary support for Machine Learning (ML) and Artificial Intelligence (AI) processes. • Evaluates emerging technologies, industry trends, and architectural patterns related to data engineering, analytics, artificial intelligence, and cloud platforms and provides recommendations for adoption. • Manages and optimizes Microsoft Fabric platform performance by monitoring capacity utilization, query performance, Lakehouse and Warehouse workloads, Delta table maintenance (OPTIMIZE/VACUUM), data partitioning, storage consumption, and pipeline execution to ensure scalable, reliable, and cost-effective data operations. • Develops and maintains monitoring, alerting, and observability solutions for Microsoft Fabric pipelines, notebooks, Lakehouses, Warehouses, semantic models, and related data platform components to ensure operational reliability and rapid issue resolution. • Performs other duties as assigned.
• Collaborate with cross-functional teams to understand data and analytical requirements and translate them into effective Microsoft Fabric–based data solutions. • Design, develop and maintain end-to-end data ingestion and transformation processes using Microsoft Fabric Data Pipelines, Dataflows Gen2, and Notebooks. • Format, cleanse, and store data in a structured manner to ensure data quality and accessibility for reporting and analysis, following Medallion Architecture • Work closely with reporting and analytics teams to ensure datasets are accurate, well-modeled, and optimized for Power BI and other analytical workloads. • Create comprehensive developer documentation for data pipelines, processes, and solutions. • Participate in data governance activities to ensure data integrity, security, and compliance. • Utilize source control to manage and track changes to data pipelines and codebase. • Provide level II/III support to troubleshoot and resolve data-related issues and inquiries. • Provide ancillary support for Machine learning (ML) and Artificial Intelligence (AI) processes • Stay updated with industry trends and best practices related to data warehousing, ETL, and data engineering.


