Overview
UltaHost is a fast-growing global web hosting and cloud infrastructure company delivering high-performance, reliable, and scalable technology solutions to customers worldwide. We are now looking for an AI Systems Architect to help design and build the technical foundation of this next stage.
- Design and build infrastructure for running AI and LLM workloads on GPU-enabled environments.
- Deploy, benchmark, optimize, and operate self-hosted and third-party LLMs.
- Build production AI applications, agents, copilots, APIs, RAG systems, and automation.
- Connect AI systems with UltaHost infrastructure, products, customer portals, support systems, databases, and internal tools.
- Explore how technologies such as GPUs, Proxmox, containers, open-source LLMs, vector databases, and modern AI frameworks can become part of UltaHost's technology stack.
- Identify internal processes where AI and deterministic automation can reduce repetitive work and improve efficiency.
- Build reusable AI infrastructure and services that can support multiple future applications instead of isolated one-off experiments.
- Help UltaHost evolve toward an AI-enabled hosting and cloud platform.
- Design, deploy, configure, and operate infrastructure for LLM and AI workloads using GPU-enabled servers.
- Work with GPU environments and understand practical considerations related to: GPU compute, VRAM capacity, model size, precision, quantization, concurrency, batching, utilization, throughput, latency, thermal and power constraints, and workload allocation.
- Evaluate the hardware and infrastructure requirements of different AI models and use cases.
- Design virtualization strategies for AI workloads using Proxmox VE, KVM, virtual machines, Linux containers, and Docker.
- Configure or support GPU and PCIe passthrough in virtualized environments where appropriate.
- Design secure resource-isolation models for internal and potentially customer-facing AI workloads.
- Contribute to GPU infrastructure capacity planning, availability, monitoring, backup, and disaster-recovery strategies.
- Develop repeatable deployment processes rather than relying on manual server configuration.
- Work with Infrastructure and Engineering teams to create stable environments for development, testing, staging, and production AI workloads.
- Design and maintain Proxmox-based infrastructure for AI and general-purpose workloads.
- Work with Proxmox clusters, virtual machines, LXC containers, networking, storage, resource allocation, templates, and backups.
- Design infrastructure with appropriate isolation, high availability, performance, and operational simplicity.
- Automate VM, container, and infrastructure provisioning using APIs, infrastructure-as-code, scripts, or orchestration tools.
- Help standardize deployment patterns across GPU servers and AI application environments.
- Troubleshoot performance, networking, storage, virtualization, and hardware-resource issues.
- Evaluate when workloads should run on bare metal, virtual machines, containers, or orchestration platforms.
- Deploy and operate open-source LLMs on private/self-hosted infrastructure.
- Evaluate and work with inference technologies such as vLLM, Ollama, LiteLLM, Hugging Face, TGI, or comparable platforms.
- Benchmark models based on quality, latency, throughput, VRAM consumption, concurrency, and operating cost.
- Understand and apply inference optimization techniques such as quantization, batching, caching, context management, and model routing where appropriate.
- Compare self-hosted models with commercial APIs such as OpenAI, Anthropic Claude, and Google Gemini and select the appropriate architecture for each use case.
- Design model gateways and reusable inference APIs that can serve multiple UltaHost applications.
- Build resilient model integrations with appropriate fallbacks, retries, rate limits, timeouts, and error handling.
- Design and build production-ready AI applications rather than demonstration-only chat interfaces.
- Develop internal AI assistants connected to authorized company knowledge, documentation, support content, operational runbooks, product information, and other approved data sources.
- Build RAG systems using embeddings, vector databases, metadata filtering, reranking, access controls, and retrieval evaluation.
- Develop AI agents capable of using controlled tools, APIs, databases, and internal services.
- Build multi-step AI workflows and agent orchestration where the complexity is justified by the use case.
- Implement structured outputs, schema validation, tool calling, memory, state management, context management, and execution controls.
- Create internal AI applications for technical and non-technical teams.
- Build customer-facing AI applications or AI-enabled hosting products where strategically relevant.
- Distinguish between tasks that require probabilistic model reasoning and tasks that should be implemented through deterministic software or traditional automation.
- Avoid unrestricted LLM-generated infrastructure actions and use controlled, validated tool interfaces instead.
- Design appropriate human-approval mechanisms for sensitive, unusual, or high-risk AI actions.
- Build reusable AI services, APIs, and platform components that can support multiple products and departments.
- Integrate AI systems with existing UltaHost platforms, portals, databases, billing systems, support systems, monitoring platforms, and internal applications.
- Work with REST APIs, webhooks, databases, queues, caches, scheduled jobs, and event-driven workflows.
- Work with relational and vector databases such as PostgreSQL/pgvector, Qdrant, or comparable technologies.
- Build backend services using Python, TypeScript/Node.js, or other appropriate technologies.
- Work with Engineering to integrate AI functionality into existing and future customer-facing products.
- Work with Engineering teams to integrate AI functionality into existing and future UltaHost products.
- Ensure that prototypes intended for production are converted into properly tested, documented, and maintainable software.
- Automate infrastructure provisioning, configuration, deployment, and maintenance.
- Work with tools such as: Terraform, Ansible, GitLab CI/CD, GitHub Actions, n8n, scripts, internal APIs, or comparable automation technologies.
- Create idempotent and repeatable infrastructure workflows where possible.
- Identify repetitive operational tasks that can be eliminated or reduced through deterministic automation or AI-assisted workflows.
- Build automations that connect infrastructure, engineering, support, and business systems.
- Evaluate whether a business problem requires AI, standard software development, workflow automation, or a hybrid solution.
- Reduce unnecessary manual dependencies without introducing unsafe or difficult-to-maintain automation.
- Monitor AI infrastructure and applications for performance degradation, model failures, infrastructure issues, and abnormal behavior.
- Apply least-privilege principles across infrastructure, applications, services, and agent tools.
- Design secure handling for credentials, secrets, API keys, service accounts, SSH access, and model-provider access.
- Implement authorization boundaries for AI applications and infrastructure actions.
- Consider tenant isolation when AI systems interact with customer data or customer infrastructure.
- Ensure that one customer, user, application, or workload cannot access another customer’s data or resources.
- Design controlled execution mechanisms for agents interacting with servers, APIs, or sensitive systems.
- Implement structured and restricted tool interfaces rather than passing unrestricted model-generated shell commands directly to production environments.
- Establish appropriate: approval gates; action-risk levels, allowlists, deny lists, validation rules, audit trails, and rollback controls.
- Address LLM-specific risks including: prompt injection, indirect prompt injection, hallucination, data leakage, unauthorized tool use, unsafe action generation, and excessive permissions.
- Ensure that AI-generated recommendations and actions are traceable and auditable where required.
- Work with departments such as Support, Sales, Marketing, Finance, HR, Operations, and Engineering to identify repetitive or inefficient workflows.
- Determine whether each problem is best solved through AI, traditional automation, software development, or a combination.
- Build AI agents, scripts, workflows, dashboards, and internal tools that reduce manual work.
- Integrate AI capabilities into existing business workflows rather than creating disconnected AI demos.
- Evaluate new AI technologies, APIs, open-source projects, models, and platforms and recommend technologies worth adopting.
- Document solutions and help teams understand how to use AI systems effectively.
- Strong Linux administration and troubleshooting skills, preferably with Ubuntu, Debian, or comparable distributions.
- Hands-on experience with KVM-based virtualization or equivalent enterprise virtualization technologies.
- Practical experience with Proxmox VE or closely related virtualization platforms.
- Understanding of virtual machines, Linux containers, virtual networking, storage, resource isolation, templates, snapshots, and backups.
- Strong Docker and containerization knowledge.
- Understanding of networking fundamentals including TCP/IP, DNS, ports, routing, firewalls, reverse proxies, TLS, and service connectivity.
- Experience operating, debugging, or supporting production infrastructure.
- Practical understanding of GPU-based AI workloads.
- Familiarity with GPU environments and the relationship between models, GPU compute, and VRAM.
- Understanding of how model size, precision, quantization, context length, batching, and concurrency affect infrastructure requirements.
- Ability to estimate and evaluate infrastructure requirements for different LLM and AI workloads.
- Familiarity with GPU drivers, runtimes, containers, passthrough, and monitoring concepts.
- Ability to investigate GPU utilization and inference-performance issues.
- Strong practical understanding of LLMs, Generative AI, RAG, embeddings, vector search, tool calling, AI agents, context management, prompts and AI workflows.
- Hands-on experience building and deploying real AI-powered applications or systems.
- Experience working with commercial LLM APIs, AI development platforms such as Lovable, Replit, Cursor, Bolt, Make, Zapier, n8n, LangChain, OpenAI API, Claude API or equivalent platforms.
- Experience with open-source LLM ecosystems and a practical understanding of self-hosted model deployment.
- Understanding of LLM evaluation, hallucination management, structured output validation, latency, reliability, and cost optimization.
- Ability to decide when an LLM is appropriate and when deterministic software or automation is the better engineering solution.
- Strong Linux fundamentals, preferably Ubuntu/Debian or comparable distributions.
- Hands-on experience with virtualization and/or containerized environments.
- Practical knowledge of Docker and modern deployment practices.
- Experience with Proxmox, KVM, VMware, Hyper-V, Kubernetes, or comparable infrastructure technologies.
- Understanding of networking fundamentals including DNS, TCP/IP, firewalls, ports, routing, and service connectivity.
- Experience designing, deploying, debugging, or operating production infrastructure.
- Good understanding of SaaS platforms, hosting businesses, customer portals, admin panels, CRM systems, or sales funnels.
- Strong scripting or programming experience with Python and/or JavaScript/TypeScript/Node.js.
- Experience integrating APIs, webhooks, databases, backend services, and third-party systems.
- Familiarity with PostgreSQL, MySQL/MariaDB, Redis, vector databases, or comparable technologies.
- Experience designing reliable automation and backend workflows.
- Working knowledge of Git and modern software development/deployment practices.
- Ability to create clean UI/UX designs or work closely with designers to produce modern, high-quality interfaces.
- Experience with monitoring, centralized logging, alerting, health checks, and production troubleshooting.
- Understanding of high availability, failure recovery, backups, and disaster recovery principles.
- Strong debugging and root-cause-analysis skills.
- Security mindset with an understanding of least privilege, credential management, isolation, and controlled system access.
- Good communication skills and ability to explain AI ideas to technical and non-technical teams.
- Ability to research, test, compare, and implement new AI tools quickly.
- Strong analytical and problem-solving ability.
- Ability to take ownership of a technical problem from initial investigation through architecture, implementation, testing, and production deployment.
- Ability to communicate technical decisions clearly with both technical and non-technical stakeholders.
- Professional working proficiency in English.
- Ability to work independently in a fully remote environment.
- At least 5 years of hands-on experience in one or more of the following areas: infrastructure engineering, platform engineering, DevOps, site reliability engineering, cloud engineering, backend systems, virtualization, or production systems administration.
- At least 2–3 years of practical experience building, deploying, operating, or integrating AI-powered applications, LLM systems, or AI infrastructure.
- Demonstrable experience taking technical solutions from initial concept or prototype into a production environment.
- Experience owning technical problems across architecture, implementation, deployment, monitoring, and troubleshooting.
- Hosting, cloud infrastructure, datacenter, VPS, dedicated-server, or SaaS environments.
- Proxmox VE and Proxmox Backup Server.
- Proxmox clustering, templates, backups, API automation, and operational troubleshooting.
- KVM virtualization and PCIe/GPU passthrough.
- ZFS, Ceph, shared storage, or distributed storage environments.
- NVIDIA GPU infrastructure and CUDA ecosystem familiarity.
- vLLM, Ollama, LiteLLM, Text Generation Inference, NVIDIA Triton, or similar technologies.
- LLM quantization approaches such as AWQ, GPTQ, GGUF, or comparable methods.
- LoRA, adapter-based fine-tuning, or model customization.
- Multi-GPU inference or distributed AI workloads.
- Model gateways and multi-model routing.
- LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, or comparable orchestration frameworks.
- Model Context Protocol and tool-based AI integrations.
- Qdrant, pgvector, Weaviate, Milvus, Pinecone, or similar vector technologies.
- Terraform and Ansible.
- Kubernetes or container-orchestration platforms.
- GitLab CI/CD, GitHub Actions, or comparable CI/CD systems.
- n8n, Make, or comparable workflow-automation platforms.
- Redis, message queues, and asynchronous job-processing systems.
- Prometheus, Grafana, OpenTelemetry, Sentry, or comparable observability platforms.
- Vault or comparable secrets-management platforms.
- SSO, identity, and access-management systems.
- WHMCS, cPanel/WHM, HostBill, Blesta, customer portals, billing systems, or hosting automation.
- Incident response, infrastructure hardening, and production-security practices.
- Building internal developer platforms or self-service infrastructure.
- Creating customer-facing infrastructure or platform products.
- Experience exposing AI capabilities through reusable internal or public APIs.
- Linux administration
- KVM-based virtualization
- Proxmox VE
- Docker
- Containerization
- GPU-based AI workloads
- LLM Engineering
- Generative AI
- RAG
- Embeddings
- Vector search
- Tool calling
- AI agents
- Python
- JavaScript/TypeScript/Node.js
- PostgreSQL
- MySQL/MariaDB
- Redis
- Git
- Monitoring and troubleshooting
- Ubuntu
- Debian
- Proxmox VE
- KVM
- Docker
- vLLM
- Ollama
- LiteLLM
- Hugging Face
- TGI
- PostgreSQL
- pgvector
- Qdrant
- Python
- TypeScript
- Node.js
- Terraform
- Ansible
- GitLab CI/CD
- GitHub Actions
- n8n
✨ Our intelligent job search engine discovered this job and republished it for your convenience.
Please be aware that the job information may be incorrect or incomplete. The job announcement remains the property of its original publisher. To view the original job and its full details, please visit the job's URL on the owner’s page.
Please clearly mention that you have heard of this job opportunity on https://ijob.am.

