A startup wants to add an AI-powered virtual assistant to their mobile app. They have limited in-house AI expertise and need a solution that can be integrated quickly with minimal infrastructure management. Which deployment pattern is MOST suitable?
Trap 1: Implement an asynchronous processing queue for all user requests
An asynchronous queue decouples request handling from processing; it provides no natural-language model, hosting, or assistant integration. It suits bursty, latency-tolerant workloads needing backpressure. The scenario demands a managed prebuilt conversational service, not a messaging pattern.
Trap 2: Train and deploy a custom model on an on-premises server
On-premises training and hosting demands GPU hardware, MLOps, and model maintenance the startup cannot staff, and delays integration. It is the right choice when data residency or strict control forbids third-party hosting. The stem instead requires minimal infrastructure management and fast delivery.
Trap 3: Deploy the model on edge devices for offline inference
Edge deployment requires exporting, optimising, and shipping models to devices, plus update and telemetry pipelines — heavy engineering for a team with no AI expertise. It is correct when offline inference and low latency are mandatory. Here connectivity and rapid managed integration are the priorities.
- A
Implement an asynchronous processing queue for all user requests
Why it fails: An asynchronous queue decouples request handling from processing; it provides no natural-language model, hosting, or assistant integration. It suits bursty, latency-tolerant workloads needing backpressure. The scenario demands a managed prebuilt conversational service, not a messaging pattern.
- B
Train and deploy a custom model on an on-premises server
Why it fails: On-premises training and hosting demands GPU hardware, MLOps, and model maintenance the startup cannot staff, and delays integration. It is the right choice when data residency or strict control forbids third-party hosting. The stem instead requires minimal infrastructure management and fast delivery.
- C
Deploy the model on edge devices for offline inference
Why it fails: Edge deployment requires exporting, optimising, and shipping models to devices, plus update and telemetry pipelines — heavy engineering for a team with no AI expertise. It is correct when offline inference and low latency are mandatory. Here connectivity and rapid managed integration are the priorities.
- D
Use a cloud-based AI microservice (e.g., Amazon Lex, Azure Bot Service) with a pre-built model
A cloud-based AI microservice with a pre-built model removes the need to train, host or scale models, directly satisfying the limited in-house expertise and minimal infrastructure management constraints. Integration is largely API configuration, enabling rapid deployment within the mobile app.