LM Studio
US - New York - New York
View Company Profile /
<< Go Back
The role involves developing and optimizing the inference stack for on-device and cloud environments, including integrating new engines and multimodal models. The engineer will also improve performance across various runtimes and contribute to open-source projects.
Requirements: Candidates need significant experience building production ML systems or inference runtimes with strong proficiency in Python and C++. Deep knowledge of transformer architectures and experience profiling CPU/GPU workloads are essential.
Key Skills: Python, C++, PyTorch, Transformer Architectures, GPU Profiling, CPU Profiling, Llama.cpp, MLX, ExecuTorch, VLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan, ROCm
Benefits: Competitive salary, Equity grants, Medical healthcare, Vision healthcare, Dental healthcare, Catered team lunch, Expensed dinners, Flexible PTO, Flexible WFH
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX