Syllo logo
Syllo

The litigation workspace for the AI era

Staff Software Engineer, AI Inference

AI EngineerMachine Learning EngineerFull TimeRemoteLeadTeam 51-200Since 2019Company SiteLinkedIn

Location

United States

Posted

1 day ago

Salary

$190K - $230K / year

Seniority

Lead

Bachelor DegreeEnglishDistributed SystemsPythonRustGo

Job Description

Staff Software Engineer, AI Inference

Syllo

• Lead the design and development of our production inference platform. • Define the technical roadmap for inference infrastructure, model serving, and runtime optimization. • Build and operate scalable, cost-effective systems for serving large language models in production. • Evaluate and integrate modern inference technologies, frameworks, and serving runtimes. • Optimize latency, throughput, GPU utilization, memory efficiency, and infrastructure cost. • Develop systems for model deployment, traffic routing, autoscaling, scheduling, observability, and operational excellence. • Partner with ML engineers to productionize new models and inference techniques. • Establish benchmarking methodologies to evaluate new models, runtimes, and hardware. • Make key architectural decisions around when to build internally versus leverage open-source or commercial solutions. • Mentor engineers as the team grows and help establish engineering best practices for AI infrastructure.

Job Requirements

  • Significant experience designing and operating production AI inference systems.
  • Experience building or leading production LLM serving infrastructure.
  • Deep experience with one or more modern inference runtimes and frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, or comparable technologies.
  • Strong background in distributed systems, backend infrastructure, or high-performance platform engineering.
  • Experience optimizing inference performance across GPU workloads, including latency, throughput, batching, memory utilization, and serving efficiency.
  • Experience operating GPU infrastructure in production.
  • Strong proficiency in Python and at least one systems programming language (such as Go, Rust, or C++).
  • Proven ability to lead technical architecture for complex infrastructure initiatives.
  • Excellent communication skills and the ability to influence technical direction across engineering teams.

Benefits

  • health insurance

Related Job Pages

More AI Engineer Jobs

Full TimeRemoteTeam 1,001-5,000Since 2004

• Design and implement an enterprise AI platform with LLM-based routing, intelligent model selection, fallback strategies, and cost optimization • Build and maintain MCP (Model Context Protocol) gateway infrastructure enabling seamless integration of enterprise tools and internal data with AI agents • Architect LangGraph-based orchestration frameworks for complex agentic workflows, enabling developers to compose multi-step AI applications • Implement RAG (Retrieval-Augmented Generation) platform components: vector search integration, prompt optimization, and chunking strategies • Establish comprehensive observability using Langfuse and similar tools: tracing, cost monitoring, token usage analytics, and LLM quality metrics. • Design and optimize APIs for LLM interactions, enabling developers to access platform capabilities seamlessly. • Build deployment and orchestration pipelines using Kubernetes (K8s) with ArgoCD for GitOps-driven AI service deployments. • Establish CI/CD infrastructure using GitLab CI/CD for automated testing, deployment, and versioning of AI platform components. • Implement security, authentication, and rate limiting for AI endpoints protecting against abuse and ensuring multi-tenant isolation. • Create self-service developer interfaces and documentation enabling non-AI-experts to build sophisticated agentic applications. • Design prompt management and versioning systems, enabling teams to collaborate on prompts and track performance across versions. • Lead technical initiatives, conduct code reviews, mentor engineers, and foster AI-first engineering culture.

Poland
zł20.4K - zł25.9K / month
Eastern Bank logo

AI Engineer

Eastern Bank

Based in Boston, we're a community-minded bank serving customers for over 200 years. #JoinUsForGood

AI Engineer1 day ago
Full TimeRemoteTeam 1,001-5,000Since 1818H1B No Sponsor

• Build, deploy, and maintain AI agents and automated workflows using platforms such as Claude, Microsoft Copilot, Azure AI Services, Power Platform, and SharePoint. • Create scalable retrieval-augmented generation (RAG) pipelines and prompt-engineering frameworks. • Prototype and productionize AI solutions that improve day-to-day operations across the bank. • Collaborate with cross-functional teams to understand business processes, pain points, and opportunities for automation. • Translate non-technical requirements into clear technical specifications and solution designs. • Serve as a trusted technical advisor supporting embedded AI initiatives across business divisions. • Develop and optimize retrieval systems, prompt architectures, and agent workflows to improve accuracy and effectiveness. • Evaluate and refine AI solutions based on user feedback, performance metrics, and business outcomes. • Assist in establishing governance processes for enterprise AI tools, including access management, usage policies, and documentation. • Promote responsible AI practices, ensuring compliance with data privacy, security, and regulatory requirements. • Support administration and enablement of AI platforms across the organization.

Massachusetts
$72.2K - $110K / year

Role Description As a Junior Software Developer (AI & Agentic Systems), you will design and implement cutting-edge agentic workflows and AI-driven applications. This role sits at the intersection of software engineering, artificial intelligence, and distributed systems—focused on building intelligent automation that empowers users to orchestrate complex tasks seamlessly. You will work closely with product owners, UI/UX designers, and other developers to build and optimize systems that enable autonomous and semi-autonomous agents to interact with data, APIs, and users. The work we do is diverse, challenging, and rewarding. Agility PR Solutions develops state-of-the-art tools that help public relations professionals discover media influencers and derive actionable insights from global media coverage. In this role, you will contribute to both: - Backend systems (Java, big data platforms like Hadoop/Solr) - Modern AI application layers (TypeScript, agent frameworks, LLM integrations) You will solve problems related to large-scale data processing, distributed workloads, and intelligent orchestration, while also contributing to evolving AI-driven product capabilities. At Agility PR Solutions, we value collaboration, curiosity, and continuous learning. You’ll be part of a team that supports growth, knowledge sharing, and innovation. Qualifications - Degree in Computer Science or a related field - Hands-on experience with Java development and REST APIs - Working knowledge of TypeScript / JavaScript - Familiarity with AI/ML integrations, including: - Large Language Models (LLMs) - Agentic frameworks (e.g., LangChain, LangGraph) - Strong problem-solving skills and willingness to learn new technologies - Experience with SQL - Experience with Linux - Experience with Git - Experience with Maven Requirements - Nice to have: - Experience with agent orchestration patterns or workflow engines - Exposure to prompt engineering and evaluation techniques - Understanding of distributed systems or big data technologies (Hadoop, Solr) Benefits - Fully remote work environment - Collaborative culture – and key tools enabling it - Competitive compensation package - Health, Dental & Vision benefits - RRSP matching - Life Insurance - Employee Assistance Program (EAP) - Career Development & Progression opportunities - Paid Vacation, Personal Days and Sick days - Flex Fridays in Summer, Week off between Christmas and New Years' - No Internal Meetings Fridays Compensation for this role is expected to fall within the range of $65,000 - 75,000 annually. The final offer will reflect each candidate’s experience, skills, and internal equity.

Worldwide
C$65K - C$75K / year
Tiger Analytics logo

Enterprise AI Architect

Tiger Analytics

AI & Analytics for today’s business challenges.

AI Engineer1 day ago
Full TimeRemoteTeam 1,001-5,000Since 2011H1B Sponsor

• Define enterprise AI architecture, standards, and technology roadmap. • Design and deliver production-ready Generative AI and Agentic AI solutions. • Architect scalable AI platforms leveraging cloud-native technologies and modern data ecosystems. • Lead AI solution design, technical reviews, and architecture governance. • Partner with engineering, data, security, and business teams to deliver enterprise AI initiatives. • Evaluate emerging AI technologies and recommend architecture best practices. • Ensure AI solutions meet security, governance, compliance, and performance requirements. • Mentor engineering teams and drive AI adoption across the organization.

Texas