More on Temu
IT y programación
Machine Learning Engineer - Inference Optimization - vacante remota
More on Temu
Featherless AI
El enlace lleva al anuncio original. Donator no recibe candidaturas.
Nuevas vacantes en Telegram
Cada vacante remota nueva, en cuanto se publica.
@donator_jobs_es
Machine Learning Engineer - Inference Optimization - IT y programación, remoto
La empresa Featherless AI busca cubrir el puesto de Machine Learning Engineer - Inference Optimization. La vacante pertenece al área de IT y programación y es totalmente remota. La empresa no pone ninguna restricción sobre el lugar de residencia, así que usted puede postular desde donde vive.
El anuncio destaca Machine Learning, así que la experiencia con esa herramienta concreta es lo que va a decidir.
La empresa no publicó ninguna cifra, eso se aclara en la entrevista. Sobre el horario, el anuncio no pone ninguna condición.
Esta vacante pasó una revisión automática: los anuncios que exigen un permiso de trabajo extranjero, patrocinio de visa, una nacionalidad concreta o residencia en un país determinado no entran en la lista.
Herramientas
En resumen
- Empresa
- Featherless AI
- Área
- IT y programación
- Quién puede postular
- Desde cualquier país del mundo
- Modalidad
- Totalmente remoto
- Herramientas
- Machine Learning
- Tipo de contrato
- Tiempo completo
- Publicada
- 24 de septiembre de 2026 (ayer)
- Activa
- hasta el 23 de noviembre de 2026
- Fuente
- Himalayas
Vacantes nuevas por correo: IT y programación
El tablón se actualiza varias veces al día. Escribimos solo cuando aparece una vacante nueva ya verificada. Nada de spam y no hace falta cuenta.
¿Ya está suscrito? Gestionar sus preferencias
Descripción de la empresa
About the Role We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale . You’ll work at the intersection of research and production-turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users. This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains. What You’ll Do
* Optimize inference latency, throughput, and cost for large-scale ML models in production
* Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO)
* Implement and tune techniques such as:
* Quantization (fp16, bf16, int8, fp8)
* KV-cache optimization & reuse
* Speculative decoding, batching, and streaming
* Model pruning or architectural simplifications for inference
* Collaborate with research engineers to productionize new model architectures
* Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks)
* Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups
* Improve system reliability, observability, and cost efficiency under real workloads
What We’re Looking For
* Strong experience in ML inference optimization or high-performance ML systems
* Solid understanding of deep learning internals (attention, memory layout, compute graphs)
* Hands-on experience with PyTorch (or similar) and model deployment
* Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations)
* Experience scaling inference for real users (not just research benchmarks)
* Comfortable working in fast-moving startup environments with ownership and ambiguity
Nice to Have
* Experience with LLM or long-context model inference
* Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton)
* Experience optimizing across different hardware vendors
* Open-source contributions in ML systems or inference tooling
* Background in distributed systems or low-latency services
Why Join Us
* Real ownership over performance-critical systems
* Direct impact on product reliability and unit economics
* Close collaboration with research, infra, and product
* Competitive compensation + meaningful equity at Series A
* A team that cares about engineering quality, not hype
Originally posted on Himalayas
El texto se conserva en el idioma original de la empresa, porque en ese mismo idioma se postulará usted.
Preguntas frecuentes sobre esta vacante
¿Puedo postularme al puesto de Machine Learning Engineer - Inference Optimization desde donde vivo?
Sí. Para esta vacante, Featherless AI acepta candidatos de cualquier país del mundo, así que usted no necesita un permiso de trabajo de otro país. El anuncio pasó una revisión automática: si la empresa hubiera exigido un permiso de trabajo, patrocinio de visa o residencia en un país determinado, no estaría en este tablón.
¿Qué remuneración se indica?
Para esta vacante, Featherless AI no publicó ninguna cifra. La mayoría de los anuncios remotos no publica un monto: se acuerda en la entrevista.
¿Cómo me postulo?
Usted se postula directamente ante la empresa, a través del anuncio original publicado en Himalayas. Donator no recibe candidaturas, no cobra comisión y no guarda ningún CV.
¿Qué tipo de vacante es esta?
Un puesto totalmente remoto en el área de IT y programación. Los anuncios híbridos y todo lo que exija presencia en la oficina no se publican en este tablón.
Certificados gratuitos para esta vacante
Este anuncio pide Machine Learning. Abajo están las credenciales gratuitas que cubren justo esas herramientas.
OCI AI Foundations Associate
exam is free tooOracle · ~10 h
A free certification on the AI track, exam included, taken inside MyLearn. An Agentic AI Foundations version also exists.
strong brandValid: 18-24 monthsAI Fundamentals and Generative AI Essentials
digital badgeIBM SkillsBuild · ~12 h
Credly badges an employer recognises. The tests require 80% to pass.
strong brandAnthropic Academy: 21 courses
certificateAnthropic · ~20 h
Courses on the Claude API, MCP and Claude Code. Each carries a free certificate, and Georgia is a supported country.
moderate weightDeep Reinforcement Learning course
certificateHugging Face · ~30 h
Entirely free with no deadline. 80% earns a Certificate of Completion, 100% a Certificate of Honors.
moderate weight
Vacantes similares
Application Security Engineer (AppSec)
$30 - $100 por horamicro1 · ayer
NuevaDesde cualquier lugarNode.jsPythonAI Researcher - Multilingual Data
Featherless AI · ayer
NuevaDesde cualquier lugarPythonMachine LearningSenior Software Engineer - React / Node.js
Adalo · ayer
NuevaDesde cualquier lugarReactNode.jsEngineering Manager - MLOps & Analytics
Canonical · ayer
NuevaDesde cualquier lugarPythonMachine Learning