donator JOBS

More on Temu

IT and programming

Machine Learning Engineer - Inference Optimization - remote job

Featherless AI

From anywhereRemote
Apply

The link takes you to the original advert. Donator takes no applications.

New vacancies on Telegram

Every new remote vacancy, posted as it lands.

@donator_jobs_en

Open the channel
Posted: (today)Active until: until 23 November 2026

Machine Learning Engineer - Inference Optimization - IT and programming, remote

Featherless AI is hiring a Machine Learning Engineer - Inference Optimization. The role sits in IT and programming and is fully remote. The company places no restriction on where the candidate lives, so you can apply from anywhere.

The advert singles out Machine Learning, so experience with that one tool is what will decide it.

The employer did not publish a figure; that is settled at interview. The advert sets no condition on working hours.

This vacancy passed an automated check: the list excludes any advert requiring a foreign work permit, visa sponsorship, a particular citizenship, or residence in a specific country.

In brief

Company
Featherless AI
Category
IT and programming
Who may apply
From anywhere in the world
Work mode
Fully remote
Tools
Machine Learning
Employment type
Full time
Posted
24 September 2026 (today)
Active
until 23 November 2026
Source
Himalayas

Get new IT and programming jobs by email

The board updates several times a day. We write only when a new, checked vacancy appears. No spam, no account needed.

Already subscribed? Manage your settings

The employer's description

About the Role We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale . You’ll work at the intersection of research and production-turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users. This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains. What You’ll Do

* Optimize inference latency, throughput, and cost for large-scale ML models in production

* Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO)

* Implement and tune techniques such as:

* Quantization (fp16, bf16, int8, fp8)

* KV-cache optimization & reuse

* Speculative decoding, batching, and streaming

* Model pruning or architectural simplifications for inference

* Collaborate with research engineers to productionize new model architectures

* Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks)

* Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups

* Improve system reliability, observability, and cost efficiency under real workloads

What We’re Looking For

* Strong experience in ML inference optimization or high-performance ML systems

* Solid understanding of deep learning internals (attention, memory layout, compute graphs)

* Hands-on experience with PyTorch (or similar) and model deployment

* Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations)

* Experience scaling inference for real users (not just research benchmarks)

* Comfortable working in fast-moving startup environments with ownership and ambiguity

Nice to Have

* Experience with LLM or long-context model inference

* Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton)

* Experience optimizing across different hardware vendors

* Open-source contributions in ML systems or inference tooling

* Background in distributed systems or low-latency services

Why Join Us

* Real ownership over performance-critical systems

* Direct impact on product reliability and unit economics

* Close collaboration with research, infra, and product

* Competitive compensation + meaningful equity at Series A

* A team that cares about engineering quality, not hype

Originally posted on Himalayas

The text is kept in the employer's original language, because that is the language you will apply in.

This vacancy was published on Himalayas. Original advert (Himalayas)

Frequently asked questions about this job

Can I apply for Machine Learning Engineer - Inference Optimization from where I live?

Yes. Featherless AI accepts candidates for this vacancy from anywhere in the world, so you need no work permit for a foreign country. The advert passed an automated check: had the employer required a work permit, visa sponsorship or residence in a specific country, it would not be on this board.

What pay is stated?

Featherless AI did not publish pay for this vacancy. Most remote adverts publish no figure; it is settled at interview.

How do I apply?

You apply directly to the employer, through the original advert published on Himalayas. Donator takes no applications, charges no commission and stores no CV.

What kind of vacancy is this?

A fully remote role in IT and programming. Hybrid adverts and anything requiring office attendance are not published on this board.

Free certificates for this vacancy

This advert asks for Machine Learning. Below are the free credentials that cover exactly those tools.

All free certificates: IT and cloud

Similar jobs

Picked on Temu

Remote job: Machine Learning Engineer - Inferen… | Donator