Senior Software Engineer, Block Production
Location
California + 1 moreAll locations: California | New York
Posted
9 hours ago
Salary
$180K - $300K / year
Seniority
Senior
Job Description
Senior Software Engineer, Block Production
Anza
• Develop and Optimize Block Production: Design, implement, and optimize the mechanisms for block production to enhance the throughput and stability of the Solana network. • Efficient Scheduling: Develop and refine scheduling mechanisms for both block production and block verification, enabling efficient transaction processing, fair resource allocation, and predictable performance. • Ensure Security and Integrity: Identify and mitigate potential security vulnerabilities within the pipeline, ensuring robust protection against emerging threats. • Scalability and Performance: Work on improving the scalability of the pipeline to handle increasing transaction volumes and validator participation. • Testing and Validation: Create and execute comprehensive tests to validate the reliability and efficiency of the block production, block verification, and scheduling mechanisms. • Collaboration: Collaborate with cross-functional teams to ensure the seamless integration and functioning of the block production components. • Documentation and Code Review: Maintain thorough documentation and conduct peer code reviews to uphold high standards of code quality and consistency. • Shared-Memory Systems: Design and maintain high-performance shared-memory interfaces for efficient data movement across the validator pipeline.
Job Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
- 3+ years of hands-on experience with core infrastructure software and distributed systems
- Strong proficiency in systems programming languages such as Rust, C or C++
- Experience with distributed systems and blockchain technology
- Ability to analyze complex systems and identify potential issues
- Knowledge of common security threats and best practices in securing block production processes
- Experience with performance profiling and optimization techniques.
- Excellent teamwork and communication skills.
Benefits
- A dynamic, fast-paced environment
- Direct impact on the security and scalability of blockchain technology
- Opportunities for innovation and problem-solving.
Related Guides
Related Categories
Related Job Pages
More Production Engineer Jobs
• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps. • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner. • Make reliability and on-call performance easier for the company to understand and improve over time.
• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Coordinate the right response • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners.
• Respond to live incidents • Maintain accurate incident records and handover • Improve the reliability system • Coordinate the right response • Protect developers from routine pages • Deliver timely and accurate impact assessments • Use AI throughout triage and diagnosis • Produce a clear incident record • Validate that customer outcome has recovered • Keep ownership, uncertainty, decisions and next actions visible
• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps.

