Senior Data Engineer

Data EngineerData EngineerFull TimeRemoteSeniorTeam 201-500

Location

United States

Posted

1 day ago

Salary

0

Seniority

Senior

Job Description

Senior Data Engineer

Knowledge Management, Inc.

Role Description The Senior Data Engineer designs, builds, and maintains the data architecture and pipelines that support SBA OIG's Technology Solutions Division (TSD) in its loan fraud detection and investigative mission. The role works within SBA's Microsoft Azure cloud environment to migrate, transform, and structure data assets so that TSD's Data Analytics team can perform agile inquiries and systematic machine learning in support of audits and investigations. - Provide highly skilled and authoritative expertise on data engineering methods and best practices, including code-first development approaches and modern pipeline design patterns. - Design, implement, and maintain an efficient, secure, stable, and flexible data architecture, with all assets managed via source control. - Design, implement, and maintain ELT/ETL pipelines for processing source data in Azure Synapse and Azure Machine Learning. - Review, maintain, and improve existing architecture and pipelines, including periodic audits addressing bottlenecks, deprecated dependencies, and architecture drift. - Establish quality controls for pipeline maintenance, and introduce error handling, logging mechanisms, and validation checks. - Incorporate source control for all pipelines and data analytics codebases to enable iterative code development while maintaining data architecture stability. - Optimize the ingestion, processing, and storage of a wide variety of datasets and data types, including modern columnar formats such as Parquet. - Develop self-service capabilities for SBA OIG analysts to query and export data for investigations and audits. - Coordinate with the Senior Data Scientists to ensure the architecture supports machine learning algorithms and data pipelines in Azure Machine Learning. - Develop standard operating protocols governing the authoring, development, validation, publishing, execution, and monitoring of all data pipelines and assets in the Azure environment. - Provide detailed documentation of the data architecture, including data dictionaries, entity-relationship diagrams, and pipeline process maps. - Maintain and expand the environment with additional datasets and services upon request, following a defined intake and testing process prior to production deployment. - Stay current with emerging AI tools relevant to data engineering and contribute to exploratory efforts evaluating automation and large language model-assisted capabilities. Qualifications - Five (5) years of hands-on experience maintaining SQL databases and conducting advanced operations in SQL and T-SQL. - Five (5) years of hands-on experience designing, implementing, and maintaining ELT/ETL processes in cloud-based data analytics environments. - Three (3) years of hands-on experience working in Azure Synapse and Azure Machine Learning within the modern data stack; DP-203 or an equivalent certification is preferred. - Three (3) years of hands-on experience manipulating data in Python, with required proficiency in Pandas; PySpark or Polars experience is preferred, along with experience developing reusable, modular code. - Preferred: implementing pipelines and infrastructure using code-first approaches, including Python SDK, CLI, REST APIs, or infrastructure-as-code tooling. - Preferred: implementing source control and CI/CD workflows. - Preferred: demonstrated familiarity with AI coding assistants and large language model integration patterns. Requirements - A Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field satisfies the education standard set forth in PWS Section 5.2.1. - In the absence of a bachelor's degree, five (5) years of applied work experience in data engineering, computer science, data science, machine learning, mathematics, or a related field will satisfy this requirement. - Microsoft Certified: Azure Data Engineer Associate, or an equivalent Azure data engineering certification, is preferred consistent with the Azure Synapse and Azure Machine Learning experience. - Certifications supporting source control, CI/CD, or infrastructure-as-code practices are preferred but not required. Benefits - Health, dental, and vision insurance - 401(k) retirement plan - Paid time off (PTO) and holidays - Group Term Life and Accidental Death and Dismemberment Insurance - Voluntary Term Life Insurance - Short and Long-term disability insurance

Related Categories

Related Job Pages

More Data Engineer Jobs

CodeRoad logo

Data Engineer

CodeRoad

Kurs JavaScript Online!

Data Engineer1 day ago
Full TimeRemoteTeam 1-10H1B No Sponsor

Role Description The Data Engineer must be comfortable working virtually as part of one or more customer engagements. You will work with stakeholders to assist with data-related technical issues and support their data infrastructure needs. Position Location: Remote (Latin America). Time Zone Requirements: This team operates on US East/West Coast time zones. Candidates must be willing to adjust schedules to meet specific project needs. How you’ll make an impact: - Pipeline Architecture: Create and maintain optimal data pipeline architecture. - Data Assembly: Assemble large, complex data sets that meet functional and non-functional business requirements. - Process Improvement: Identify, design, and implement internal process improvements: automating manual processes, optimizing data delivery, and re-designing infrastructure for greater scalability. - Infrastructure: Build the infrastructure required for optimal extraction, transformation, and loading (ETL) of data from a wide variety of sources using SQL and Cloud technologies (Azure/GCP). - Tooling: Create data tools for analytics and data scientist team members that assist them in building and optimizing our product into an innovative industry leader. Qualifications - Advanced English (B2-C1, spoken and written). - 4+ years as a Data Engineer experience building processes supporting data transformation, data structures, metadata, dependency, and workload management. - Experience with Databricks, Python, and SQL (Scala is a plus). - Familiarity with Apache Spark, Data Factory, Synapse, or BigQuery. - Experience building Big Data pipelines via Streams and/or Batches. - Strong skills related to working with unstructured datasets and performing root cause analysis to identify opportunities for improvement. Benefits - 100% Remote - Holidays Off - Paid Time Off - Health insurance assistance program - Competitive Pay (USD) - Excellent teamwork and work environment - Training

Latin America (LATAM)
eSimplicity logo

Senior Data Engineer

eSimplicity

An engineering firm that delivers high-quality Healthcare IT, Cybersecurity, and Telecommunication solutions.

Data Engineer1 day ago
Full TimeRemoteTeam 51-200Since 2016H1B No Sponsor

Role Description Identifies and owns all technical solution requirements in developing enterprise-wide data architecture. - Creates project-specific technical design, product and vendor selection, application, and technical architectures. - Provides subject matter expertise on data and data pipeline architecture and leads the decision process to identify the best options. - Serves as the owner of complex data architectures, with an eye toward constant reengineering and refactoring to ensure the simplest and most elegant system possible to accomplish the desired need. - Ensures strategic alignment of technical design and architecture to meet business growth and direction and stays on top of emerging technologies. - Develops and manages product roadmaps, backlogs, and measurable success criteria and writes user stories. - Responsible for expanding and optimizing our data and data pipeline architecture, as well as optimizing data flow and collection for cross-functional teams. - Supports software developers, database architects, data analysts, and data scientists on data initiatives and ensures that the optimal data delivery architecture is consistent throughout ongoing projects. - Creates new pipeline development and maintains existing pipeline; updates Extract, Transfer, Load (ETL) process; creates new ETL feature development; builds PoCs with Redshift Spectrum, Databricks, etc. - Implements, with the support of project data specialists, large dataset engineering: data augmentation, data quality analysis, data analytics (anomalies and trends), data profiling, data algorithms, and develops data strategy recommendations. - Assembles large, complex data sets that meet non-functional and functional business requirements. - Identifies, designs, and implements internal process improvements, including re-designing data infrastructure for greater scalability, optimizing data delivery, and automating manual processes. - Builds required infrastructure for optimal extraction, transformation, and loading of data from various data sources using AWS and SQL technologies. - Builds analytical tools to utilize the data pipeline, providing actionable insight into key business performance metrics, including operational efficiency and customer acquisition. - Works with stakeholders, including data, design, product, and government stakeholders, and assists them with data-related technical issues. - Writes unit and integration tests for all data processing code. - Works with DevOps engineers on CI, CD, and IaC. - Reads specs and translates them into code and design documents. - Performs code reviews and develops processes for improving code quality. Qualifications - All candidates must pass public trust clearance through the U.S. Federal Government. - Bachelor’s degree in Computer Science, Engineering, or a related technical field; OR in lieu of a degree, 10 additional years of relevant professional experience and 8 years of specialized experience may be substituted. - 10+ years of total professional experience in the technology or data engineering field. - Extensive Data pipeline experience using Python, Java, and cloud technologies. - Expert data pipeline builder and data wrangler who enjoys optimizing data systems and building them from the ground up. - Self-sufficient and comfortable supporting the data needs of multiple teams, systems, and products. - Experienced in designing data architecture for shared services, scalability, and performance. - Experienced in designing data services including API, metadata, and data catalog. - Experienced in data governance process to ingest (batch, stream), curate, and share data with upstream and downstream data users. - Ability to build and optimize data sets, ‘big data’ data pipelines, and architecture. - Ability to perform root cause analysis on external and internal processes and data to identify opportunities for improvement and answer questions. - Excellent analytic skills associated with working on unstructured datasets. - Ability to build processes that support data transformation, workload management, data structures, dependency, and metadata. - Demonstrated understanding and experience using software and tools, including big data tools like Kafka, Spark, and Hadoop; relational NoSQL and SQL databases including Cassandra and Postgres; workflow management and pipeline tools such as Airflow, Luigi, and Azkaban; AWS cloud services including Redshift, RDS, EMR, and EC2; stream-processing systems like Spark-Streaming and Storm; and object function/object-oriented scripting languages including Scala, C++, Java, and Python. - Flexible and willing to accept a change in priorities as necessary. - Ability to work in a fast-paced, team-oriented environment. - Experience with Agile methodology, using test-driven development. - Experience with Atlassian Jira/Confluence. - Excellent command of written and spoken English. - Ability to obtain and maintain a Public Trust; residing in the United States. Requirements - Federal Government contracting work experience. - Google’s Certified Professional-Data-Engineer certification, IBM Certified Data Engineer – Big Data certification, CCP Data Engineer for Cloudera. - Centers for Medicare and Medicaid Services (CMS) or Health Care Industry experience. - Experience with healthcare quality data, including Medicaid and CHIP provider data, beneficiary data, claims data, and quality measure data. - Experience with Medicaid Data. - PySpark/Spark experience. - Experience with ETL operations. - Experience supporting CMS. - Experience with AWS, MWAA (Airflow), EMR, Standard services: S3, EC2, Secrets Manager, SNS, CloudWatch, Version Control/Github, Python, and Java. Benefits - Comprehensive benefits package, including medical, dental, and vision coverage. - 401(k) retirement benefits. - Paid time off and paid holidays. - Life and disability insurance. - Additional wellness and employee support programs. - Eligibility may vary based on employment status and applicable plan terms. Reasonable Accommodation eSimplicity is committed to providing reasonable accommodations to qualified individuals with disabilities during the application and hiring process. Applicants who need assistance or an accommodation should contact Human Resources. Equal Employment Opportunity eSimplicity is an Equal Opportunity Employer, including disability and protected veteran status. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran status, disability, or any other legally protected status.

United States
$132.4K - $155K / year
Hotel Engine logo

Director, Data Engineering

Hotel Engine

Innovating business travel with a free-to-use hotel booking platform.

Data Engineer1 day ago
Full TimeRemoteTeam 201-500Since 2018H1B No Sponsor

• Strategic Vision & Roadmap: Define the 2–3 year vision for the data platform. Align data engineering initiatives with overarching company goals to ensure data is a primary driver of business growth. • Organizational Leadership: Build and scale a multi-tiered engineering organization. Foster a culture of technical excellence, mentorship, and high accountability, ensuring a robust leadership pipeline within your team. • Data Governance & Compliance: Establish and enforce global standards for data quality, security, and privacy (GDPR/CCPA). Own the framework for data cataloging and lineage to ensure "one source of truth" across the enterprise. • Operational Excellence & ROI: Optimize the total cost of ownership (TCO) for our data infrastructure. Manage vendor relationships (Snowflake, etc.) and budgets while ensuring the stack scales efficiently with company growth. • Executive Stakeholder Management: Serve as the primary liaison between Engineering and the C-suite. Translate technical debt and infrastructure needs into business value, and turn business requirements into scalable technical architectures. • Innovation & Architecture: Drive the adoption of cutting-edge technologies (AI/ML integration, real-time streaming) to keep Engine at the forefront of the industry.

United States
$224.0K - $310K / year
Boldr logo

Seasonal Data Services Associate

Boldr

Helping Companies Build Global Teams Through Ethical Outsourcing

Data Engineer1 day ago
Full TimeRemoteTeam 501-1,000H1B No Sponsor

Role Description As a Seasonal Data Services Associate, you will be working with different tools and databases to review data. You are responsible for executing processes as defined by the client and/or management. The role requires keen attention to detail while maintaining productivity at defined proficiency levels and an aptitude to learn new processes. Qualifications - Bachelor's / College Degree in any field you’re passionate about! - Previous experience in a related field is a plus. - Basic knowledge of Salesforce is a plus. - Experience in using CRM and other similar applications or tools. - Basic knowledge of cloud-based applications such as Google Drive, Google Sheets, Google Docs and MS Office applications. Requirements - Execute processes assigned by the client and/or management. - Ensure defined productivity and quality targets are met. - Data Entry (encoding handwritten details from photo forms). - Identify process improvement opportunities as they arise. - Reports process deficiencies to ensure accurate execution of tasks. - Creating and finishing projects assigned by the client. - Collaborate with internal and external teams in completing projects. - Review submissions. Benefits - Own work device - Work From Home - Training & Development

Worldwide