Lead ML Engineer - LLM Inference & NLP
Social Discovery Group
Job description
About the role
Social Discovery Group (SDG) is seeking a Lead ML Engineer to own LLM inference and NLP initiatives. You will accelerate production deployment of large language models and guide technical direction for both NLP and computer‑vision teams.
Key responsibilities
- Speed up and scale LLM inference in production using SGLang, KV and prefix caching, batching, quantization, and speculative decoding.
- Run distributed inference for models up to 1 trillion+ parameters across multi‑GPU and multi‑node setups.
- Benchmark new GPU servers, integrate them into production, and adapt serving code.
- Lead the NLP and CV teams technically: review experiments, set direction, and intervene early when issues arise.
- Train and fine‑tune language models, improve agent harnesses and chat algorithms.
- Track cutting‑edge research and open‑source work in inference and post‑training, turning insights into the ML roadmap.
- Collaborate with validation, content, and dataset preparation teams to design experiments and measure model quality.
Required profile
- Deep hands‑on experience optimizing LLM inference in production with SGLang, vLLM, or TensorRT‑LLM.
- Experience with distributed inference or training of large models (MoE, tensor/expert/pipeline parallelism, multi‑node GPU clusters).
- Strong understanding of fast inference techniques: KV cache, attention kernels, batching, quantization, GPU profiling.
- Experience training and fine‑tuning LLMs, including post‑training methods such as RLHF or DPO.
- Proven technical leadership: mentoring engineers, guiding reviews, and making architectural decisions while still coding.
- Proficiency with PyTorch, transformers, and related libraries.
- Background in AI‑focused startups or similar companies (e.g., Character AI, OpenAI) is a strong plus.
- Backend engineering experience (Python, Go, C#) and knowledge of scalable deployment systems.
- Advanced English or Russian language skills.
Required skills
- SGLang
- vLLM
- TensorRT‑LLM
- PyTorch
- Transformers
- Python
- Go
- C#
- CUDA
- Triton
- GPU profiling
- KV cache
- Attention kernels
- Batching
- Quantization
- Distributed inference
- Mixture of Experts (MoE)
- Tensor parallelism
- Pipeline parallelism
- Multi‑node GPU clusters
What we offer
- Remote full‑time opportunity.
- 28 vacation days per year plus 7 wellness days.
- Referral bonuses up to $5,000.
- 50% reimbursement for professional training, conferences, and meetings.
- Corporate discount for English lessons.
- Health benefit reimbursement up to $1,000 per year.
- Equipment allowance up to $1,000 every three years for home or co‑working spaces.
- Internal gamified gratitude system with bonuses, merchandise, and team‑building activities.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Serbia.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 day ago
Expires 1 month from now
6 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Social Discovery Group
Related job offers
-
CRM Sales and Service Functional Consultant
Bosch Srbija Belgrade -
Principal IAM Engineer – Directory Services
Syneos Health Belgrade -
Technical Project Manager – Serbian & Russian (C1+)
Life Data Lab, LLC -
Service Desk Agent L1 – Customer Support
NCR Voyix Belgrade -
Full Stack Graduate Developer – Insurance Tech
SAP Fioneer Belgrade