More on Temu
Autres offres
AI Researcher - Inference Optimization - poste à distance
More on Temu
Featherless AI
Le lien mène à l'annonce d'origine. Donator ne reçoit aucune candidature.
AI Researcher - Inference Optimization - Autres offres, à distance
L'entreprise Featherless AI recrute : AI Researcher - Inference Optimization. Le poste relève de la catégorie Autres offres et se fait entièrement à distance. Aucune restriction n'est posée sur le lieu de résidence du candidat, vous pouvez donc postuler de partout.
Les principaux outils cités dans l'annonce sont les suivants : Python, Machine Learning. Votre CV doit montrer des exemples concrets de ces compétences.
L'entreprise n'a publié aucun chiffre, cela se règle en entretien. L'annonce ne pose aucune condition d'horaires.
Cette offre a passé un contrôle automatique : les annonces qui exigent un permis de travail étranger, un parrainage de visa, une nationalité précise ou une résidence dans un pays nommé n'entrent pas dans la liste.
Outils
En bref
- Entreprise
- Featherless AI
- Catégorie
- Autres offres
- Qui peut postuler
- De n'importe quel pays du monde
- Mode de travail
- Entièrement à distance
- Outils
- Python, Machine Learning
- Type de contrat
- Temps plein
- Publiée
- 25 septembre 2026 (il y a 4 jours)
- Validité
- jusqu'au 24 novembre 2026
- Origine
- Himalayas
Nouvelles offres par e-mail : Autres offres
Le tableau est actualisé plusieurs fois par jour. Nous n'écrivons que lorsqu'une nouvelle offre vérifiée paraît. Pas de spam, aucun compte à créer.
Déjà abonné ? Gérer vos préférences
La description de l'entreprise
Role Overview We are seeking an AI Researcher with deep experience in inference optimization to design, evaluate, and deploy high-performance inference systems for large-scale machine learning models. You will work at the intersection of model architecture, systems engineering, and hardware-aware optimization , improving latency, throughput, and cost efficiency across real-world production environments. Key Responsibilities
* Research and develop techniques to optimize inference performance for large neural networks.
* Improve latency, throughput, memory efficiency, and cost per inference .
* Design and evaluate model-level optimizations (quantization, pruning, KV-cache optimization, architecture-aware simplifications).
* Implement systems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization).
* Benchmark inference workloads across hardware accelerators.
* Collaborate with engineering teams to deploy optimized inference pipelines .
* Translate research insights into production-ready improvements .
Required Qualifications
* Strong background in machine learning, deep learning, or AI systems .
* Hands-on experience optimizing inference for large-scale models .
* Proficiency in Python and modern ML frameworks (e.g., PyTorch).
* Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime).
* Ability to design experiments and communicate results clearly.
Preferred / Nice-to-Have Qualifications
* Experience deploying production inference systems at scale .
* Familiarity with distributed and multi-GPU inference .
* Experience contributing to open-source ML or inference frameworks .
* Authorship or co-authorship of peer-reviewed research papers in machine learning, systems, or related fields.
* Experience working close to hardware (CUDA, ROCm, profiling tools).
What Success Looks Like
* Measurable gains in latency, throughput, and cost efficiency .
* Optimized inference systems running reliably in production.
* Research ideas successfully translated into deployable systems.
* Clear benchmarks and documentation that inform product decisions.
Relevant Research Areas (Bonus)
* Long-context inference optimization
* Speculative decoding
* KV-cache compression and paging
* Efficient decoding strategies
* Hardware-aware inference design
Originally posted on Himalayas
Le texte est conservé dans la langue d'origine de l'entreprise, car c'est dans cette langue que vous postulerez.
Questions fréquentes sur cette offre
Puis-je postuler au poste de AI Researcher - Inference Optimization depuis là où je vis ?
Oui. Pour cette offre, Featherless AI accepte des candidats de n'importe quel pays du monde, vous n'avez donc besoin d'aucun permis de travail pour un autre pays. L'annonce a passé un contrôle automatique : si l'entreprise avait exigé un permis de travail, un parrainage de visa ou une résidence dans un pays précis, elle ne figurerait pas sur ce tableau.
Quelle rémunération est indiquée ?
Pour cette offre, Featherless AI n'a publié aucune rémunération. La plupart des annonces à distance ne donnent aucun chiffre, cela se règle en entretien.
Comment postuler ?
Vous postulez directement auprès de l'entreprise, via l'annonce d'origine publiée sur Himalayas. Donator ne reçoit aucune candidature, ne prend aucune commission et ne conserve aucun CV.
De quel type d'offre s'agit-il ?
Un poste entièrement à distance dans la catégorie Autres offres. Les annonces hybrides et tout ce qui impose une présence au bureau ne sont pas publiés sur ce tableau.
Certificats gratuits pour cette offre
Cette annonce demande Python, Machine Learning. Vous trouverez ci-dessous les attestations gratuites qui couvrent précisément ces outils.
Deep Reinforcement Learning course
certificateHugging Face · ~30 h
Entirely free with no deadline. 80% earns a Certificate of Completion, 100% a Certificate of Honors.
moderate weightAI Agents course
certificateHugging Face · ~25 h
The most current free course on building agents. The old 2025 deadline has been lifted.
moderate weightKaggle Learn: 17 micro-courses
certificateKaggle · ~4 h
Intro to ML, Pandas, SQL, Deep Learning, Computer Vision and more. Only a Google account is needed, and the certificate has a public link.
moderate weightOCI AI Foundations Associate
exam is free tooOracle · ~10 h
A free certification on the AI track, exam included, taken inside MyLearn. An Agentic AI Foundations version also exists.
strong brandValid: 18-24 months