RemoteListed Jul 14, 2026
LLM Inference Engineer
NEAR AI
- AI-Infrastructure
- GPU-Computing
- Inference-Optimization
- LLM-Inference-Engineering
- Machine-Learning-Infrastructure
- Senior-AI-Inference-Engineer
- Senior-AI-LLM-Engineer
Assessed from original listing evidence
The role
Job description
Locations: San Francisco or Remote
About The Role
The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.
What You'll Be Doing
- Architect and maintain production high-traffic LLM serving systems.
- Optimize throughput, latency, and cost for leading open-source LLMs.
What We're Looking For
- Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
- Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
- Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
- Strong problem-solving skills and ability to communicate technical ideas clearly.
We'd Love If You Have
- Experience with Trusted Execution Environments (TEE).
- Active contributor to open-source LLM inference engines.
Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.
Originally posted on Himalayas
Keep exploring
Related remote jobs
Senior Software Engineer FullStack (AI Apps)
NEAR AI
Machine LearningAI Engineer — AWS Bedrock AgentCore & Production Agentic Systems | LATAM & Europe
Toptal
Machine LearningAI-Native Software Developer
OnTheGoSystems
Machine LearningSr. Product Manager, Enterprise AI
Trueblue
Machine LearningAI Researcher — AI Architecture Research
Featherless Ai
Machine Learning