Remote roles optimizing large-language-model serving — throughput, latency, and cost. Think vLLM, TensorRT, quantization, and GPU efficiency. Inferr surfaces engineers with verified production inference work.
There are no open llm inference engineer roles on the board at this moment. Build your verified profile now so you're the first match when one opens — or browse every remote AI role.
They make model serving fast and cheap — optimizing throughput and latency with tools like vLLM and TensorRT, applying quantization, and tuning GPU utilization for production traffic.