Vicobi AI
A FastAPI backend integrating speech recognition and receipt processing, turning pretrained AI models into a deployable API.

Overview
Vicobi AI is a backend for a personal finance application, centered on two AI workflows: converting speech into structured data and extracting information from receipt images. The project was developed during an internship at AWS.
Problem
Transaction information may arrive as spoken input or receipt images, while the backend needs structured data that can be stored and reused. The pretrained models required a unified integration layer to accept files, process their outputs, and expose the results through an API.
Role
I researched and developed the speech and receipt AI workflows, then integrated them into a FastAPI backend. My scope covered preparing the inference pipelines, connecting the processing stages, and defining APIs for application use.
Solution
- Used PhoWhisper to transcribe Vietnamese audio before extracting structured transaction data.
- Combined image processing, EasyOCR, and a PyTorch classifier to read and validate receipts.
- Used AWS Bedrock to normalize information extracted from speech and receipt content.
- Exposed the workflows through FastAPI and persisted their results in MongoDB.
- Packaged the service with MongoDB and Qdrant using Docker Compose.
Technical decisions
Instead of training new models, the project reused pretrained models and focused on the inference pipeline and backend integration. A service layer separates AI logic from the FastAPI routers so that each workflow can evolve independently. AWS Bedrock handles structured information extraction, while speech recognition, OCR, and receipt classification remain with specialized models.
Results
The project delivered a backend API connecting both AI workflows through speech- and receipt-processing endpoints. The repository contains no benchmark data, so the outcome is reported as a functional integration rather than an accuracy score or quantified improvement.
Lessons learned
- How to move pretrained models into a backend pipeline beyond model experimentation.
- How to separate API, service, and AI layers to reduce coupling between application logic and models.
- How to coordinate file conversion, inference, structured extraction, and persistence within one request flow.
Keep exploring