SECTION I · THE BRIEF
Brief #95901Updated 22 AUG 2026PALO ALTO, CAAshbyY COMBINATOR
Employbl Company Profile

Software Engineer, Cloud

Ollama is a software firm that develops local inference platform designed to run large language models.

Location
Palo Alto, CA
Company size
2–10
Posted
3d ago
Via
Ashby
Section II · Premium ProfileMembers only
  • 01Comp band & equity packageLocked
  • 02Seniority & experience requirementsLocked
  • 03Interview process & rubricLocked
  • 04Hiring manager & team contextLocked
  • 05Growth trajectory in this roleLocked
  • 06Offer & decision timelineLocked

7-day free trial · $25/mo · cancel anytime

Ollama logo

Software Engineer, Cloud · Ollama

View company profile
Job title
Software Engineer, Cloud
Job location
Palo Alto
Job description

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.

What you'll do

  • Build and scale the inference platform that serves every request from ollama.com.

  • Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.

  • Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.

  • Build the reliability, observability, and cost controls for our team and customers

You may be a fit if

  • You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.

  • You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.

  • You've worked with Kubernetes, GPU scheduling, or inference infrastructure.

  • You think in terms of reliability, SLOs, and honest capacity planning.

  • Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

View job listing ↗
The Saturday Briefing

Get the Saturday tech briefing

New company profiles, funding moves, and who’s hiring across the market — every Saturday morning.

Where this role is based

Palo Alto, CA

Loading map…

Ollama headquarters

Palo Alto, CA

Company size

210 employees

Founded

2021

Total raised

$80,125,000

View company profile ↗

Funding rounds

  • Series B$65M
  • Series A$15M
  • Pre Seed$125K