FriendliAI

Supercharge Generative AI Inference Efficient, fast, and reliable generative AI inference solution for production

friendli.ai

FriendliAI

Supercharge Generative AI Inference Efficient, fast, and reliable generative AI inference solution for production

friendli.ai

HQUS

Team Size49

Open Jobs2

Total Funding-

Latest FundraiseUnknown

TL;DR

What they do: Managed inference cloud for deploying and serving large language and multimodal models with performance optimizations

HQ: Redwood City, California

Founded: 2021

Recent funding: $20M seed extension (Aug 28, 2025)

CEO / Founder: Byung‑Gon (Gon) Chun

Company Overview

Problem Domain

AI inference infrastructure for large language and multimodal models

Founded

2021

Industry

Software Development

Funding Track Record

Seed extension- 2025-08-28

$20M

Round announced to expand AI inference platform, go-to-market, and product development

Investor Signal

“Capstone Partners led the $20M seed extension with participation from Sierra Ventures, Alumni Ventures, KDB, and KB Securities”

Founders

What we do

Join the Team

Solution Architect - AI Inference Specialist

On-SiteSan Francisco Bay Area, US

On-Site • San Francisco Bay Area, US

Related Companies

Company	HQ	Industry	Total Funding
FuriosaAI	🌍Undisclosed	Data and AnalyticsDeepTechInformation TechnologyManufacturing	$266M
Baseten	🇺🇸US	Software	$585M
quadric, Inc	🇺🇸Burlingame, US	Consumer ProductsDeepTechHardwareManufacturing	$74M
Modular	🇺🇸US	Data and AnalyticsDeepTechInformation TechnologySoftware	$380M
GenBio AI	🇺🇸Palo Alto, US	BiotechnologyDeepTechEducation	-

About the job

FriendliAI is seeking a Forward Deployed Engineer (FDE) to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container.

Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product.

You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position.

Key Responsibilities

Design and implement large-scale deployment architectures for LLM and multimodal inference
Deploy and manage containerized workloads across Kubernetes clusters
Diagnose production issues, such as performance bottlenecks, and implement temporary fixes as needed
Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows
Develop scripts, Helm charts, and Terraform modules that simplify repeated deployments
Contribute field insights to shape our platform reliability, observability, and scaling strategies
Lead workshops, technical sessions, or webinars to help customers master infrastructure best practices

Qualifications

3+ years of experience in cloud infrastructure, DevOps, or reliability engineering
Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
Proficiency with Kubernetes, Docker, Terraform, and Helm
Strong foundation in distributed systems, networking, and performance tuning
Experience with GPU-based computing and generative AI model serving workloads
Strong technical background in backend systems or AI tooling
Experience operating workloads on AWS, GCP, or OCI
Excellent problem-solving and debugging skills in real-world environments

Preferred Experience

Experience deploying large models (LLMs, diffusion models) on GPUs or clusters
Familiarity with inference frameworks (Triton, vLLM, TensorRT, DeepSpeed-Inference)
Familiarity with observability stacks (Prometheus, Grafana, Loki, ELK, OTEL)
Understanding of networking security and compliance frameworks (e.g., SOC 2)
Experience supporting on-prem or hybrid-cloud deployments

Benefits

A front-row seat to the generative AI infrastructure revolution
Competitive compensation and benefits package
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up and top-tier hardware support
Flexible working hours and a highly collaborative environment
We offer competitive compensation, startup equity, health insurance, and other benefits.

About us

FriendliAI is building the next-generation AI inference platform that accelerates the deployment of large language and multimodal models with unmatched performance and efficiency. Our infrastructure powers high-throughput, low-latency workloads for global organizations and integrates directly with Hugging Face, providing instant access to over 480,000 open-source models. We are on a mission to deliver the world’s best platform for AI inference.

Startup jobs. A lot of them.

Your next opportunity is in here somewhere. Sign up to explore 52,000+ startups and their open roles. No spam. No gamification. Just jobs.

52,000+

Startups

66,000+

Open Roles

1,300+

New This Week

Software Engineer

InternshipNiš, RS

Internship • Niš, RS

Data Scientist

ContractCambridge, GB

Contract • Cambridge, GB

DevOps Engineer

Part-timeNovi Sad, RS

Part-time • Novi Sad, RS

AI Researcher

ContractNew York, US

Contract • New York, US

Product Designer

InternshipJerusalem

Internship • Jerusalem

Software Engineer

Part-timeHaifa

Part-time • Haifa