Job Closed

This listing is no longer active.

Senior/Staff Data Scientist

Location

United States

Posted

28 days ago

Salary

$165K - $300K / year

Seniority

Lead

Job Description

Senior/Staff Data Scientist

BNSF Railway

Role Description Be part of a team that values safety, inclusion, and excellence. As a member of our team, you will play a role in supporting the movement of essential products and materials that help feed, clothe, supply, and power communities throughout America and the world. Are you ready to drive change? If you are passionate about making a difference and eager to advance your career in a dynamic and supportive environment, we want you on our team! Join us in reshaping the future of freight rail and discover a fulfilling career where your contributions matter. This position is open to candidates who are currently authorized to work in the United States. We are also open to sponsoring H-1B transfers, TN nonimmigrant status, and STEM OPT candidates with at least 2 years of remaining eligibility. Key responsibilities may include: - Lead cross-functional collaboration to identify and define analytic initiatives, formulate strategies, and develop solutions to achieve business goals through effective use of data machine learning models. - Apply data science skills to analyze large, complex datasets and identify meaningful patterns that lead to actionable insights and data-driven solutions to business problems. - Lead the development and deployment of advanced machine learning models to forecast outcomes and optimize workflows. - Engage closely with stakeholders to grasp their requirements and deliver actionable insights that drive strategic decision-making. - Design and present compelling visualizations and reports to effectively communicate analytical findings. - Oversee the maintenance and enhancement of data pipelines to uphold data quality and precision. - Keep abreast of emerging trends and breakthroughs in data science and machine learning fields. - Proficiently extract, aggregate, and transform data from SQL and NoSQL databases, leveraging languages like R, Python, or other relevant tools for analysis and modeling. - Build, test, and validate statistical and machine learning models and analyses using Python, R, or other appropriate language as part of overall solution development. - Lead implementation of analytic solutions into reporting platforms or production systems by leading the solution design, development, testing, and monitoring. - Demonstrate operational excellence by monitoring, troubleshooting, and resolving production issues, including participating in a 24/7 on-call rotation. Qualifications - Minimum 6 years of experience with building optimization algorithms or relevant experience. - Advanced proficiency in programming languages such as Python, R, SQL, and Java. - Demonstrated expertise in utilizing data visualization tools to communicate insights clearly and effectively. - In-depth experience with data science cloud platforms and their integration into business solutions. - Exceptional intellectual curiosity and a proven ability to thrive in a fast-paced, collaborative team environment. - Track record of rapidly acquiring new technical skills and adapting to cutting-edge technologies. - Strong communication skills to articulate technical concepts to diverse audiences with clarity and professionalism. - Deep understanding and application of statistical analysis and advanced machine learning techniques. - Must understand and have proficiency in the core architecture of LLMs (e.g., transformers and attention mechanisms). - Experience with prompt engineering techniques, including chain-of-thought prompting, Retrieval-Augmented Generation (RAG), fine-tuning of language models, and evaluation methodologies. - Experience with Vector Databases and embeddings. - Experience with Model Fine-tuning. - Experience with GPU optimization. Preferred Qualifications - Bachelor's degree or higher in Operations Research, Computer Science, Industrial Engineering or a related field. Ph.D is a plus. - Knowledge of geospatial analytics, route optimization, and GIS concepts. - Previous hands-on experience with AI/Machine Learning frameworks and tools, showcasing innovative solutions. - Extensive background in Rail, Shipping, Airline, Logistics, Warehousing, Supply Chain, or Transportation industries, or in the High-Tech sector. - Proficiency in leveraging open-source libraries and frameworks to drive data science initiatives. - Seasoned in Agile methodologies like Scrum, Kanban, or SAFe for efficient project management and delivery at a senior level. Benefits - An industry-leading 401(k) and renowned Railroad Retirement program. - A range of robust health care options for you and your dependents (including domestic partners), including medical, dental, vision, telemedicine, mental health, cancer support, and high-quality care network options. - Health care spending accounts (HSA) with employer contributions, as well as life and disability insurance, provided at no cost. - Family benefits including parental, pediatric and family building support, adoption and surrogacy reimbursement, and dependent care spending account (with employer match). - Access to discounts on travel, gym memberships, counseling services and wellness support. - Annual bonus (Incentive Compensation Program). - Generous leave / time off policies.

Related Categories

Related Job Pages

More Data Scientist Jobs

Full TimeRemoteTeam 2-10Since 2019

• Translate customer goals — such as improving differential diagnosis, evaluating a clinical note summarizer, testing a RAG-based medical literature assistant, or creating preference data for patient-facing chatbots — into dataset specifications, taxonomies, rubrics, sampling plans, and acceptance criteria. • Make multimodal health AI a core focus: design training and evaluation datasets across clinical text, medical images, waveforms, structured EHR data, claims, trial data, medical literature, patient communications, payer policies, drug information, and other clinical artifacts, as well as use cases such as clinical reasoning, medical QA, note summarization, medical coding, patient communication, utilization management, and literature synthesis. • Design evaluations for retrieval-augmented and source-grounded health AI systems, including evidence citation, faithfulness, contraindication handling, guideline adherence, source freshness, and failure modes caused by incomplete, conflicting, or stale context. • Define sampling strategies, label schemas, inter-annotator agreement targets, adjudication workflows, SME review patterns, and quality thresholds in partnership with Language Data Scientists, clinicians, biomedical experts, and quality teams. • Build statistical and ML checks that make healthcare datasets trustworthy: stratified sampling across specialties and patient subgroups, bias and representation analysis, leakage detection, distribution shift checks, uncertainty estimates, reliability metrics, and subgroup performance analysis. • Partner with Applied Research Scientists and AI/ML Research Engineers to instrument datasets into evaluation and post-training pipelines, including rubric-grounded LLM-as-judge prompts, regression suites, model comparison workflows, experiment tracking, and model-improvement feedback loops. • Evaluate health AI behavior beyond surface accuracy: calibration, hallucination on safety-critical content, refusal appropriateness, robustness under ambiguity, equity across patient subgroups, and safe handoff in agentic or workflow-integrated systems. Reason concretely about clinical workflow fit: where outputs enter care delivery, what evidence a clinician or reviewer would need to trust them, when uncertainty must be surfaced, and how patient-facing, clinician-facing, payer, pharma, and operational use cases differ in risk. • Own data quality from source intake through delivery, including de-identified clinical text, medical literature, synthetic cases, structured records, client policies, and knowledge bases, with attention to PHI/PII handling, provenance, audit trails, versioning, and compliance documentation. • Stay current on the health AI landscape — regulatory developments such as FDA guidance on AI/ML-enabled medical devices and EU AI Act health provisions, benchmark releases such as MedQA, MedMCQA, and HealthBench, and emerging clinical evaluation methodology. • Support customer discovery and proposal work by scoping dataset programs, sizing annotation and SME review effort, identifying regulatory or data-access constraints, and explaining methodology choices to client clinical and ML leadership. • Contribute to Innodata internal IP: reusable health-domain taxonomies, evaluation rubrics, golden datasets, clinical review playbooks, dataset quality checks, and methodology templates.

United States
$150K - $175K / year
Full TimeRemoteTeam 2-10Since 2019

• Translate customer goals — such as improving financial reasoning, building an eval suite for earnings-call summarization, or evaluating an AML/fraud copilot — into concrete dataset specifications, taxonomies, rubrics, and acceptance criteria. • Design training and evaluation datasets across the financial AI surface: financial QA, filings and earnings analysis, credit and underwriting, fraud/AML investigation, and compliance, among other financial workflows. • Foreground unstructured and multimodal financial data in dataset design — PDFs, scanned statements, tables, charts, and call transcripts — used by analysts, advisors, compliance reviewers, and operations teams. • Design datasets and evaluations for retrieval-augmented and source-grounded systems: evidence citation and faithfulness to source documents, data freshness, conflict resolution across sources, and failure modes caused by incomplete or incorrectly parsed context. • Evaluate agentic and workflow-integrated financial AI systems: tool use, retrieval, transaction boundaries, escalation behavior, and controls that prevent unsafe or unauthorized actions. • Develop evaluation methodology that goes beyond surface accuracy — numerical consistency, hallucination rates on high-risk claims, refusal and escalation appropriateness, robustness under ambiguity, and fairness across protected or sensitive customer segments. • Define sampling strategies, label schemas, and adjudication workflows with Language Data Scientists and finance SMEs; write annotation guidelines that make subjective finance-domain judgments explicit, calibratable, and auditable. • Build the statistical and ML tooling that makes large financial datasets trustworthy: stratified sampling across products, markets, and modalities; bias analysis; leakage detection; and distribution shift checks, among other reliability checks. • Build evaluation and dataset-quality evidence to support financial-services model risk management: assumptions, limitations, validation results, and residual risks, packaged as reproducible evidence. • Partner with the AI/ML Research Engineer to instrument datasets into training, evaluation, and monitoring pipelines — rubric-grounded LLM-as-judge prompts, regression suites, and continuous monitoring. • Own data quality end-to-end, from intake through delivery: PII handling, provenance tracking, versioning, and modality-specific QA checks. • Reason about financial workflow context: where AI outputs enter analyst, advisor, compliance, risk, or customer-facing workflows; what evidence a reviewer needs to trust them; and when uncertainty must be surfaced. • Support the Technical Solutions Architect during customer discovery and proposals: scoping dataset programs, sizing annotation effort, and explaining methodology to client stakeholders. • Stay current on the financial AI landscape: regulatory developments, benchmark releases, and emerging evaluation methodology for finance-domain models. • Contribute to Innodata internal IP: reusable taxonomies, evaluation rubrics, golden datasets, and methodology templates.

United States
$150K - $175K / year
Cross Border Talents logo

Head of Data

Cross Border Talents

🌎 Your international recruitment partner for hard to find professionals and jobs all over the globe.

Data Scientist30 days ago
Full TimeRemoteTeam 201-500Since 2013H1B No Sponsor

• Build and own the company's data function from the ground up • Audit and improve existing data infrastructure, tooling, and processes • Lead a lean data team while remaining hands-on with analysis • Leverage AI to automate workflows and increase team productivity • Analyze large datasets to answer strategic business questions • Deliver investment-grade analyses, dashboards, and executive briefings • Establish data governance, reporting standards, and analytical best practices • Partner with Product, Finance, Operations, and Leadership to drive business decisions • Continuously improve the company's AI-enabled data capabilities

Arizona
Job Closed
Cross Border Talents logo

Head of Data

Cross Border Talents

🌎 Your international recruitment partner for hard to find professionals and jobs all over the globe.

Data Scientist30 days ago
Full TimeRemoteTeam 201-500Since 2013H1B No Sponsor

• Build and own the company's data function from the ground up • Audit and improve existing data infrastructure, tooling, and processes • Lead a lean data team while remaining hands-on with analysis • Leverage AI to automate workflows and increase team productivity • Analyze large datasets to answer strategic business questions • Deliver investment-grade analyses, dashboards, and executive briefings • Establish data governance, reporting standards, and analytical best practices • Partner with Product, Finance, Operations, and Leadership to drive business decisions • Continuously improve the company's AI-enabled data capabilities

New York
Job Closed