Parallel Wireless logo
Parallel WirelessPosted 1 month ago

Senior/Principal Local LLM & Generative AI Platform Engineer

On-siteKfar Saba, Central District, Israel

Full TimeSenior LevelMedium

Job Summary

Own the architecture and technical roadmap for a secure, reliable local LLM platform deployed in Parallel Wireless-controlled infrastructure. Build modular inference and model-gateway layers with stable APIs, routing, and concurrency controls while partnering with engineering, product, support, IT, security, and legal teams to prioritize high-value use cases. Design and operate RAG and enterprise-search pipelines for approved repositories, wikis, tickets, and logs, enforcing strict source-system permissions to prevent unauthorized access. Establish versioned evaluation datasets and automated offline/online testing for retrieval quality, groundedness, and latency, creating release gates and reproducible regression tests. Implement end-to-end observability for model workflows, including traces, error monitoring, and resource utilization, alongside safe tool-calling and agent workflows with sandboxing and human approval gates. Integrate the platform into existing developer environments, CI workflows, and knowledge systems through reusable SDKs and APIs, building operational foundations for production use including CI/CD, configuration registries, backups, and disaster recovery. Protect proprietary and personal information through network isolation, encryption, redaction, and defenses against prompt injection and supply-chain risks.

Required Qualifications

  • BSc or MSc in Computer Science, Computer Engineering, Electrical Engineering, Data Science, or a related field, or equivalent practical experience
  • Typically 7+ years of hands-on experience in production software, ML platform, search, data, or infrastructure engineering, including meaningful recent experience shipping LLM-powered systems
  • Strong Python engineering skills and experience designing maintainable APIs, services, libraries, and data pipelines
  • Experience with Go, Java, or C/C++
  • Strong understanding of transformer-based language models and production inference, including tokenization, context management, batching, KV caching, parallelism, quantization, structured output, tool calling, and common model failure modes
  • Demonstrated experience building production RAG or enterprise-search systems using embeddings, vector and/or lexical search, metadata filtering, reranking, source attribution, and systematic retrieval evaluation
  • Experience defining task-specific LLM evaluations using representative datasets, strong baselines, domain-expert review, automated metrics, human feedback, error analysis, and regression thresholds
  • Experience deploying and operating containerized services on Linux using Docker and Kubernetes or an equivalent orchestration environment
  • Practical experience with GPU-backed model serving, performance profiling, capacity planning, monitoring, and reliability engineering
  • Strong knowledge of distributed-system fundamentals, authentication and authorization, API security, secrets handling, encryption, auditability, and data lifecycle controls
  • Experience with Git, automated testing, CI/CD, infrastructure as code, observability, and production incident response
  • Sound technical judgment about quality, security, maintainability, hardware efficiency, and total cost—not just model benchmark scores
  • Ability to lead an ambiguous, cross-functional initiative, explain complex AI behavior in plain language, and help other teams ship safely on a shared platform

Desired Qualifications

  • Experience operating LLMs in on-premises, private-cloud, restricted-network, or air-gapped environments
  • Hands-on experience with current inference runtimes and serving systems such as vLLM, SGLang, TensorRT-LLM, llama.cpp, Ray Serve, KServe, Triton, or equivalent technologies
  • Experience optimizing inference on NVIDIA and/or AMD GPUs using CUDA, ROCm, profiling tools, tensor parallelism, pipeline parallelism, speculative decoding, prefix/KV caching, or related techniques
  • Experience with model and experiment registries, LLM tracing and evaluation platforms, vector databases, hybrid-search engines, and production data-orchestration frameworks
  • Experience with parameter-efficient fine-tuning methods such as LoRA/QLoRA, dataset curation, synthetic-data generation, distillation, and post-training evaluation
  • Experience building code intelligence, repository-aware assistants, developer tools, or IDE and CI integrations for large C/C++ and Python codebases
  • Familiarity with Active Directory or another enterprise identity provider, fine-grained document authorization, data-loss prevention, secure software supply chains, model licensing, and AI governance
  • Experience red-teaming LLM or agent systems for prompt injection, sensitive-data disclosure, poisoned retrieval content, excessive agency, and insecure output handling
  • Knowledge of telecommunications, 3GPP, RAN/Open RAN, cloud-native network functions, or technical-support workflows
  • Experience working across heterogeneous compute platforms and making performance, energy, and TCO tradeoffs for enterprise AI workloads
  • Contributions to relevant open-source AI, search, MLOps, or infrastructure projects

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce