Senior Site Reliability Engineer, Observability

Sorry, this job was removed at 10:51 p.m. (EST) on Tuesday, Aug 13, 2024
Be an Early Applicant
Remote
175K-210K Annually
5-7 Years Experience
Cloud • Information Technology • Machine Learning
We empower creators and innovators with access to GPU resources they need to work more efficiently.
The Role

CoreWeave is a specialized cloud provider, delivering a massive scale of GPU compute resources on top of the industry’s fastest and most flexible infrastructure. CoreWeave builds cloud solutions for compute intensive use cases — VFX and rendering, machine learning and AI, batch processing, and Pixel Streaming — that are up to 35 times faster and 80% less expensive than the large, generalized public clouds. Learn more at www.coreweave.com.

The Observability Team performs a critical role in enabling CoreWeave to understand, troubleshoot, and optimize complex systems by providing comprehensive insights into their behavior and performance. This team is responsible for the development, integration, and operation of observability platforms with the ultimate objective of enabling engineers across CoreWeave to do more, better. Central to the Observability Teams mission is the operation of our observability stack which leverages CoreWeave’s deep investment in the Kubernetes ecosystem.

We are seeking a senior engineer with specialization in the observability stack who can help us execute on the mission of providing a comprehensive logging and metrics ecosystem that is deeply integrated with CoreWeave’s Kubernetes platform. Integrating logging, metrics, tracing, and monitoring tools for proactive insights into system performance. This individual will work with a team of 6-8 engineers and have the opportunity to work on the full gamut of rewarding challenges that come with the business of building a cloud in a communicative, supportive, and high-performing environment. As a member of the Observability Team you will have the opportunity to:

  • Design and implement the platform that improves visibility into how the services are performing and operating.
  • Improve the performance, security, reliability, and scalability of our observability, and related services and participate in the teams on-call rotation.
  • Assist engineers in maximizing the observability stack to gain insights into the service's functionality and operation.
  • Develop dashboards, alerts, and insights into the customer experience using Grafana-ecosystem tools such as Mimir and Loki.
  • Develop meaningful insights by analyzing the gathered data.
  • Enable and evangelize the best practices around alerting. Collaborate with teams to establish observability standards.
  • Grow, change, invest in your teammates, be invested-in, share your ideas, listen to others, be curious, have fun, and, above all, be yourself.


  • You have four or more years of experience in a software or infrastructure engineering industry.
  • You enjoy helping your colleagues achieve more with less effort.
  • You have experience operating services in production and at scale and are versed in reliability engineering concepts such as the different types of testing, progressive deployments, error budgets, the role observability, and fault-tolerant design.
  • You have experience using Kubernetes with a conceptual understanding of its major components, and/or have operated Kubernetes clusters at scale for both event-driven and stateful orchestration.
  • You’re familiar with various logging and metrics systems like Prometheus, ELK, Victoria Metrics, Thanos or Grafana. You have experience with designing and operating these systems at scale.
  • You are familiar with PromQL, any other querying language and enjoy understanding the data model for observability systems. 
  • You’re comfortable with the idea of using Go as your primary programming language.
  • You know your way around a Linux distro, shell scripting, and/or the Linux storage and networking stacks.
  • You can transform problems in elastic solutions, decompose them into achievable tasks, and socialize both to your teammates.
  • You’re excited about being part of a team of diverse perspectives and backgrounds that believe in tackling challenges, growing hand in hand, and winning together.

Our compensation reflects the cost of labor across several US geographic markets. The base pay for this position ranges from $175,000-$210,000. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience.

Successful candidates will be expected to attend onboarding training at our NJ Headquarters within their first several weeks of employment, with subsequent quarterly travel requirements of 1 week duration.

If you reside within a 30-mile radius of our New Jersey, New York, or Philadelphia offices, we're excited for you to join us at the office at least three times a week, recognizing the significance we place on fostering connections, collaboration, and creativity within our office culture. Our commitment to operating as a hybrid workplace underscores our dedication to enabling our employees to tailor their work-life balance to their individual preferences.

Why CoreWeave?

At CoreWeave, we work hard, have fun, and move fast!  We’re in an exciting stage of hyper-growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: 

  • Be Curious at your Core
  • Act like an Owner
  • Empower Employees
  • Deliver Best In-Class Client Experience 
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! 

We offer a competitive salary and benefits, including:

  • Medical, dental and vision insurance - 100% paid for the employee
  • Company paid Life Insurance 
  • Voluntary supplemental life insurance 
  • Short and long-term disability insurance 
  • Flexible Spending Account
  • Tuition Reimbursement 
  • Mental Wellness Benefits through Spring Health 
  • Family-Forming support provided by Carrot
  • Paid Parental Leave 
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our offices
  • Weekly massages in NJ office
  • A casual work environment
  • Work culture focused on innovative disruption

California Consumer Privacy Act - California applicants only


What the Team is Saying

Louis
Taylor
Anthony
Matt
The Company
Roseland, NJ
600 Employees
Hybrid Workplace
Year Founded: 2017

What We Do

CoreWeave is a specialized cloud provider, delivering a massive scale of GPU compute resources on top of the industry’s fastest and most flexible infrastructure. CoreWeave builds cloud solutions for compute intensive use cases — VFX and rendering, machine learning and AI, batch processing, and Pixel Streaming — that are up to 35 times faster and 80% less expensive than the large, generalized public clouds. Learn more at www.coreweave.com.

Why Work With Us

At CoreWeave we work hard, have fun and move fast! Today we are a small, growing team of intelligent, genuine people, that value different perspectives and approaches to solving complex problems. We foster an environment that champions collaboration and prioritizes innovative solutions. Here, you are surrounded by the best.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

CoreWeave Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: Flexible
Roseland, NJ

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account