AI Data Engineer

Location

Pakistan

Posted

2 days ago

Salary

0

Seniority

Mid Level

Job Description

AI Data Engineer

Pakistan Single Window

Role Description Build and optimize ETL pipelines for large-scale AI applications, integrating APIs, web scraping, and real-time data processing. Develop and maintain scalable AI data infrastructure using PySpark, Pandas, SQL, and cloud services (Azure, AWS, GCP). Implement retrieval-augmented generation (RAG) pipelines using FAISS, ChromaDB, and Pinecone for AI-driven insights. Deploy and monitor AI applications using FastAPI, Streamlit, Docker, for real-time performance. Work with cross-functional teams to ensure data security, compliance, and AI model reliability in production environments. Qualifications - Bachelor's or Master's degree in Computer Science, Data Engineering, Artificial Intelligence, or a related field. - At least 2-4 years of experience in data engineering, AI/ML, or cloud-based AI infrastructure. - Expertise in Python, PySpark, and SQL for data transformation and large-scale processing. - Experience with cloud platforms (AWS, Azure, GCP) for AI model deployment and data pipeline automation. - Hands-on experience with vector databases (FAISS, ChromaDB, Pinecone) for efficient data retrieval. - Proficiency in containerization and orchestration tools like Docker and Kubernetes. - Strong understanding of retrieval-augmented generation (RAG) and real-time AI model deployment. - Knowledge of API development and AI service integration using FastAPI and Streamlit/Dash. - Ability to optimize AI-driven automation processes and ensure model efficiency. - Strong analytical and problem-solving skills. - Excellent communication and collaboration abilities. Requirements - Expertise in Python, PySpark, and SQL for data transformation and large-scale processing. - Experience with cloud platforms (AWS, Azure, GCP) for AI model deployment and data pipeline automation. - Hands-on experience with vector databases (FAISS, ChromaDB, Pinecone) for efficient data retrieval. - Proficiency in containerization and orchestration tools like Docker and Kubernetes. - Strong understanding of retrieval-augmented generation (RAG) and real-time AI model deployment. - Knowledge of API development and AI service integration using FastAPI and Streamlit/Dash. - Ability to optimize AI-driven automation processes and ensure model efficiency. - Strong analytical and problem-solving skills. - Excellent communication and collaboration abilities. Company Description

Related Categories

Related Job Pages

More Data Engineer Jobs

Children's Mercy KC logo

Data Engineer

Children's Mercy KC

Children’s Mercy is in the heart of Kansas City – a metro abounding in cultural experiences, vibrant communities, and thriving businesses. This is where our patients and families live, work, and play. This is a community that has embraced our hospital and we strive to say thanks by giving back. As a leader in children’s health, we engage in meaningful programs and partnerships throughout the region so that we can improve the lives of children beyond the walls of our hospital.

Data Engineer2 days ago
Full TimeRemoteTeam 5,001-10,000

Role Description The Data Engineer is responsible for building and maintaining reliable data pipelines that support Children's Mercy Integrated Care Solutions analytics and operational needs. This role contributes to the development of the data platform, ensuring quality, accuracy, and scalability, while growing technical skills and collaborating effectively with peers. - Develop, configure, maintain, test and monitor processes in support of a reliable, scalable, and evolving data & analytics platform. - Collaborate with product owners, analysts, and architects to understand business needs and optimize dataflows. - Design and implement data pipelines, support approved data models, and continuously optimize workflows for performance, reliability, and maintainability. - Develop curated, reusable, and well-modeled data assets/products (datasets, reports, dashboards) within CMICS. - Perform thorough testing of processes to ensure continuous accuracy and reliability. - Maintain technical competency by exploring emerging technologies and engaging in continuous learning. Qualifications - Master's Degree in Computer Science, Information Systems, Data Science, Software Engineering, or similar and 1-2 years experience with cloud-native platforms, distributed systems, and modern data practices including ETL/ETL methodologies or related experience. - Bachelor's Degree in Computer Science, Information Systems, Data Science, Software Engineering, or similar and 3-5 years experience with cloud-native platforms, distributed systems, and modern data practices including ETL/ETL methodologies or related experience. Benefits - Recognized as one of the best places to work in Kansas City. - Benefits plans designed to meet the changing needs of employees and their families. - Starting pay begins at $41.57/hr, determined based on education and experience. - This position is entirely remote or work-from-home, with the requirement to live in the Kansas City metro area.

United States
$86.5K / year
Bluelight Consulting logo

Data Engineer (Senior) - ETL (Python+Snowflake)

Bluelight Consulting

Bluelight is a leading software consultancy dedicated to designing and developing innovative technology that enhances users' lives. With a steadfast commitment to delivering exceptional service to our clients, Bluelight excels in its focus on quality and customer satisfaction. Our mission is not only to create cutting-edge applications but also to foster a collaborative and enriching work environment where each team member can grow and thrive. With a presence across the United States and Central/South America, Bluelight is in an exciting phase of expansion, continually seeking exceptional talent to join its dynamic and diverse community.

Data Engineer2 days ago
Full TimeRemoteTeam 201-500

Role Description We are looking for a skilled individual to join our rapidly growing team at Bluelight. This position is ideal for someone who thrives in a fast-paced, dynamic environment where everyone's opinions and efforts are valued and appreciated. You will have the opportunity to contribute to challenging and meaningful projects, developing high-quality applications that stand out in the market. We value continuous learning, personal growth, and hard work, offering a collaborative environment that promotes professional development. If you are passionate about software development and eager to be part of a growing software consultancy, we invite you to apply and join us on this exciting journey. Qualifications - Bachelor’s degree in Computer Science, Information Technology, Data Engineering, or a related field (or equivalent professional experience). - Proven experience developing end-to-end data pipelines extracting/transforming/loading data from REST APIs, relational databases, cloud storage, and flat files. - Demonstrated hands-on experience with Snowflake (virtual warehouses, streams, tasks, stages, Snowpipe, secure data sharing, and performance optimization). - Advanced SQL development skills with the ability to write complex queries, tune performance, and optimize large-scale workloads. - Experience with dimensional modeling techniques (star schemas, fact tables, dimension tables). - Familiarity with cloud-based data ecosystems, particularly Microsoft Azure. - Managing source code and CI/CD pipelines using Git and Azure DevOps (or similar). - Strong analytical, problem-solving, and detail-oriented mindset. - Excellent verbal and written communication skills; ability to collaborate in a fast-paced environment with evolving priorities. - Knowledge of data integration best practices, data governance, and enterprise data management. Requirements - Certifications: SnowPro, Azure Data Engineer Associate, or equivalent cloud data platform certifications (preferred). - Experience with big data technologies, machine learning, data science platforms, or advanced analytics (preferred). - Familiarity with BI tools such as Power BI, Tableau, or similar visualization platforms (preferred). - Experience with Agile delivery frameworks and DevOps practices (preferred). Responsibilities - Design, develop, and maintain scalable ETL/ELT pipelines using Python (PySpark), Snowflake, and cloud-native technologies. - Build reliable, efficient, and reusable data ingestion, transformation, and loading processes. - Utilize Snowflake’s architecture to design, build, and optimize modern cloud data solutions. - Implement and manage Snowflake objects (databases, schemas, tables, views, streams, tasks, stages, stored procedures). - Leverage virtual warehouses, data sharing, time travel, and automated scaling to balance performance and cost efficiency. - Apply dimensional modeling (star schemas, facts, dimensions) to build scalable enterprise data warehouses. - Extract and ingest structured and semi-structured data from REST APIs, relational databases, SaaS apps, flat files, and cloud storage. - Develop robust ingestion frameworks. - Collaborate with data architects and stakeholders to create logical and physical data models aligned with business goals. - Contribute to modern data platform concepts (data lakes, lakehouses, data mesh architectures, enterprise data catalogs). - Support integration between Snowflake and cloud-native services (primarily Azure). - Establish and enforce data engineering standards and best practices. - Implement automated data quality controls, validation frameworks, and monitoring processes. - Support data governance initiatives and maintain adherence to organizational standards. - Ensure data security, privacy, and regulatory compliance (access controls, masking policies, industry best practices). - Monitor and optimize Snowflake workloads, ETL/ELT processes, and SQL queries to meet SLAs. - Analyze warehouse utilization and recommend performance and cost-efficiency improvements. - Monitor pipelines, diagnose performance issues, and implement long-term solutions; support production environments and incident resolution. - Maintain comprehensive documentation for pipelines, flows, transformations, data models, and operational processes. - Partner with cross-functional teams (data architects, data scientists, analysts, business stakeholders) to understand requirements and provide technical expertise. Benefits - Competitive salary and bonuses, including performance-based salary increases. - Generous paid-time-off policy. - Flexible working hours. - Work remotely. - Continuing education, training, conferences. - Company-sponsored coursework, exams, and certifications.

Latin America and the Caribbean
Shawmut Services LLC logo

Data Engineer

Shawmut Services LLC

Shawmut Services is a full-service strategic planning, organizing, and grassroots mobilization firm.

Data Engineer2 days ago
Full TimeRemoteTeam 10,001+Since 2021H1B No Sponsor

• Design, develop, and maintain scalable ETL/ELT data pipelines. • Build and optimize data architectures and warehouses (e.g., Snowflake, BigQuery Etc). • Integrate data from a wide variety of sources using APIs and web browser automations. • Collaborate with company leadership to understand data needs and implement ongoing improvements in system processes. • Monitor and ensure data quality, reliability, and performance. • Automate data workflows and implement data validation and testing processes. • Maintain up-to-date documentation for data pipelines, systems, and processes.

United States
$72K - $84K / year
Full TimeRemoteTeam 1,001-5,000Since 1972H1B No Sponsor

• Fine-tuning of SLMs tailored to specific needs • Implementation of agents and integrations as required • Implementation of multi-LLM routing • Preparation of technical and business documentation • Knowledge transfer • Conduct cost–benefit analysis.

Brazil
R$11K / month