Capital Technology Group, LLC logo
Capital Technology Group, LLC

Simple Solutions for Complex Problems

Junior Data Engineer

Data EngineerData EngineerFull TimeRemoteJuniorTeam 11-50Since 2010H1B No SponsorCompany SiteLinkedIn

Location

United States

Posted

1 day ago

Salary

$75K - $115K / year

Seniority

Junior

Job Description

Junior Data Engineer

Capital Technology Group, LLC

• Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Databricks, dbt, Apache Spark (PySpark), SQL (PostgreSQL), and Python. • Develop and optimize AWS-native data solutions leveraging services including AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch. • Build high-performance data ingestion, transformation, and orchestration workflows across structured and semi-structured data using Parquet, ORC, Avro, and Apache Iceberg. • Integrate data from enterprise and external sources, including relational and NoSQL databases such as PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL platforms. • Improve the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, and continuous optimization. • Develop and maintain automated deployment pipelines using Harness and collaborate on cloud infrastructure and platform improvements. • Support data engineering efforts powering mission-critical analytics, reporting, and decision-making across large-scale federal data environments. • Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and continuously improve engineering processes.

Job Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field
  • 5 years of professional experience in data engineering or related domains
  • Strong hands-on experience with: Databricks for large-scale data engineering, ETL/ELT development, and data transformation.
  • Strong proficiency with Dbt, SQL (PostgreSQL), Python, Java, AWS (Redshift, RDS, S3, Glue, Lambda, DMS, CloudWatch), Data Pipelines, Data Modeling, System Maintenance, Data Integration, Performance Optimization.
  • AWS native.
  • Strong analytical and problem-solving skills
  • Experience working in agile, iterative software development environments
  • Ability to quickly learn and apply new technologies and domain knowledge
  • Excellent written and verbal communication skills, with the ability to explain complex topics to diverse audiences.

Benefits

  • Remote Work
  • Competitive Compensation Package
  • Medical, Dental, and Vision
  • Life Insurance, Short/Long Term Disability
  • Employee Assistance Program
  • 401(k) with 4% matching
  • Liberal PTO vacation policy
  • Generous Annual Continuing Education
  • Annual Wellness Budget
  • Bonus Incentive Programs

Related Categories

Related Job Pages

More Data Engineer Jobs

Capital Technology Group, LLC logo

Data Engineer

Capital Technology Group, LLC

Simple Solutions for Complex Problems

Data Engineer1 day ago
Full TimeRemoteTeam 11-50Since 2010H1B No Sponsor

• Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), Databricks, dbt, SQL (PostgreSQL), and AWS Glue. • Develop and optimize AWS-native data platforms leveraging AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Lambda, Step Functions, Amazon S3, Redshift, RDS, DMS, and CloudWatch. • Build high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data using Apache Iceberg, Parquet, ORC, and Avro. • Design and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies. • Integrate enterprise and external data sources across relational and NoSQL platforms including PostgreSQL, Oracle, Redshift, GraphDB, and other NoSQL databases. • Build AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vector and OpenSearch vector indexes. • Develop cloud infrastructure using CloudFormation (Infrastructure as Code), GitHub, Harness, and enterprise CI/CD pipelines while leveraging SNS, SQS, and EventBridge for event-driven architectures. • Improve the reliability, scalability, performance, and maintainability of enterprise data platforms through monitoring, troubleshooting, automation, and continuous optimization. • Support mission-critical analytics and reporting solutions within large-scale AWS-based federal data environments, implementing solutions that comply with FedRAMP and NIST 800-53 security controls. • Lead modernization initiatives migrating legacy platforms including IBM DataStage, Hadoop, RunDeck, and shell-based workflows to cloud-native AWS services. • Mentor junior engineers through technical guidance, architecture discussions, and code reviews while promoting engineering best practices. • Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts effectively to technical and non-technical stakeholders.

United States
$110K - $140K / year
JobGet logo

Principal Data Engineer

JobGet

The go-to marketplace for the Everyday Worker. We help employers meet job seekers where they are.

Data Engineer1 day ago
Full TimeRemoteTeam 51-200Since 2019

• Lead the technical direction for how data flows, scales, and powers decisions at JobGet • Own the data architecture behind the platform • Drive architectural decisions • Champion modern data stack adoption • Ensure platform reliability • Build production-grade pipelines • Establish data modeling standards • Solve complex data integration and performance challenges • Architect and evolve streaming data infrastructure • Enable machine learning at scale • Establish data governance standards • Implement data validation frameworks • Raise the technical bar through mentoring

United States
Full TimeRemoteTeam 5,001-10,000

• Collaborates with cross-functional teams to understand data and analytical requirements and translates them into effective Microsoft Fabric–based data solutions. • Mentors and provides technical leadership to Data Engineers through code reviews, design reviews, knowledge-sharing sessions, and engineering best practices. • Establishes and maintains data engineering standards, reusable frameworks, development patterns, and operational best practices to improve consistency and scalability across the data platform. • Leads root cause analysis and resolution efforts for critical production issues impacting enterprise data pipelines, reporting solutions, and platform operations. • Reviews and provides guidance on solution designs, architecture decisions, and implementation approaches to ensure alignment with enterprise standards and best practices. • Designs, develops, and maintains end-to-end data ingestion and transformation processes using Microsoft Fabric OneLake, Lakehouses, Data Pipelines, Dataflows Gen2, and Notebooks. • Formats, cleanses, and stores data in a structured manner to ensure data quality and accessibility for reporting and analysis, following Medallion Architecture. • Works closely with reporting and analytics teams to ensure datasets are accurate, well-modeled, and optimized for Power BI and other analytical workloads. • Creates comprehensive documentation for data pipelines, processes, and solutions. • Participates in data governance activities to ensure data integrity, security, and compliance. • Utilizes source control to manage and track changes to data pipelines and codebase. • Provides level II/III support to troubleshoot and resolve data-related issues and inquiries. • Provides ancillary support for Machine Learning (ML) and Artificial Intelligence (AI) processes. • Evaluates emerging technologies, industry trends, and architectural patterns related to data engineering, analytics, artificial intelligence, and cloud platforms and provides recommendations for adoption. • Manages and optimizes Microsoft Fabric platform performance by monitoring capacity utilization, query performance, Lakehouse and Warehouse workloads, Delta table maintenance (OPTIMIZE/VACUUM), data partitioning, storage consumption, and pipeline execution to ensure scalable, reliable, and cost-effective data operations. • Develops and maintains monitoring, alerting, and observability solutions for Microsoft Fabric pipelines, notebooks, Lakehouses, Warehouses, semantic models, and related data platform components to ensure operational reliability and rapid issue resolution. • Performs other duties as assigned.

Worldwide
Full TimeRemoteTeam 5,001-10,000

• Collaborate with cross-functional teams to understand data and analytical requirements and translate them into effective Microsoft Fabric–based data solutions. • Design, develop and maintain end-to-end data ingestion and transformation processes using Microsoft Fabric Data Pipelines, Dataflows Gen2, and Notebooks. • Format, cleanse, and store data in a structured manner to ensure data quality and accessibility for reporting and analysis, following Medallion Architecture • Work closely with reporting and analytics teams to ensure datasets are accurate, well-modeled, and optimized for Power BI and other analytical workloads. • Create comprehensive developer documentation for data pipelines, processes, and solutions. • Participate in data governance activities to ensure data integrity, security, and compliance. • Utilize source control to manage and track changes to data pipelines and codebase. • Provide level II/III support to troubleshoot and resolve data-related issues and inquiries. • Provide ancillary support for Machine learning (ML) and Artificial Intelligence (AI) processes • Stay updated with industry trends and best practices related to data warehousing, ETL, and data engineering.

United States