Scaling Through AI Optimization
Fetch, the popular rewards and consumer engagement platform, has successfully overhauled its machine learning pipeline to handle the immense pressure of processing millions of receipts weekly. Faced with the challenge of managing over 80 million receipt scans per week—reaching peak traffic of hundreds of scans per second—the company required a more robust, scalable infrastructure. By migrating its models to Amazon SageMaker and incorporating Hugging Face's advanced machine learning tools, Fetch has managed to cut processing latency by half, ensuring that rewards are delivered to users almost instantaneously.
The transition, which spanned a 12-month development cycle, involved the deployment of over five new machine learning models. This optimization was not merely a speed improvement; it represented a fundamental shift in how the company extracts and structures data from consumer transactions. The resulting improvements have empowered the platform to provide deeper insights to brand partners, ultimately fueling growth from 10 million to 18 million monthly active users.
Technical Synergy: SageMaker and Hugging Face
The core of this technical transformation lies in the seamless integration between Amazon SageMaker and the Hugging Face ecosystem. Fetch utilized AWS Deep Learning Containers to deploy transformer models, allowing for a simplified, managed environment that minimized manual overhead. The use of the Amazon SageMaker Hugging Face Inference Toolkit provided the company with an open-source standard for serving its complex models with high reliability.
Hardware acceleration played a critical role in the project’s success. By leveraging multi-GPU instances within Amazon SageMaker, the engineering team at Fetch was able to handle intensive data processing workloads that were previously bottlenecks. Furthermore, tools like the Amazon SageMaker Inference Recommender automated the load testing and tuning of models, which significantly reduced the time required to bring new features from the lab into the production mobile app.
Why it Matters
- Latency Reduction: By optimizing their inference pipeline, Fetch reduced the latency for their slowest scans by 50%, providing a snappier, more reliable user experience.
- Accuracy Gains: The integration of advanced model tuning and training capabilities resulted in a 200% improvement in the accuracy of their document-understanding models.
- Operational Efficiency: With standardized deployments and automated shadow testing, the team lowered the barrier for model deployment, enabling even new team members to push updates safely and effectively.
- Scalability: The architecture allows for parallel model training and seamless scaling for inference, providing a future-proof foundation for expanding into new domains like fraud prevention.
Future-Proofing Through Innovation
The success of the initiative has encouraged the Fetch team to continue pushing the boundaries of what their machine learning models can achieve. By maintaining a tech stack that utilizes the latest SageMaker features, the company is effectively future-proofing its platform. Plans are already underway to apply these sophisticated ML pipelines to broader business challenges, including enhanced fraud detection and more granular consumer behavior analysis. As Fetch continues its trajectory, the partnership between its internal engineering prowess and the scalable infrastructure of AWS stands as a blueprint for companies looking to turn vast amounts of raw consumer data into precise, actionable intelligence.











