Loading use case index…
Loading use case index…
AI use case
Hong Kong-based fintech Apoidea Group, creator of the SuperAcc document-processing platform, worked with AWS to fine-tune the Qwen2-VL-7B-Instruct vision-language model with LLa…
Core facts from this catalog record. Primary narrative lives in the hero above; full raw fields follow in the next section.
Every column from the source row, in stable order. URLs open in a new tab.
Title
How Apoidea Group enhances visual information extraction from banking documents with multimodal models using LLaMA-Factory on Amazon SageMaker HyperPod | Artificial Intelligence
Content
Hong Kong-based fintech Apoidea Group worked with AWS to fine-tune the open-source Qwen2-VL-7B-Instruct vision-language model with LLaMA-Factory on Amazon SageMaker HyperPod, raising its TEDS table-structure-recognition accuracy from 23.4 to 81.1—surpassing Anthropic's Claude 3 Haiku on the FinTabNet benchmark. The post is co-written with Apoidea Group's Ken Tsui (VP of Machine Learning), Edward Tsoi (Senior Data Scientist) and Mickey Yip (VP of Product). It frames the work as 'a transformative leap forward from traditional multistage and single-modality methods, offering an end-to-end solution for modern document processing challenges,' without a single directly quoted executive. Banking document processing—KYC, loan applications, financial spreading and multi-bank-statement reviews—has historically required large human teams to extract data from unstructured documents. Apoidea's flagship product, SuperAcc, was already deployed at 10+ financial-services clients before this engagement; the new model targets the table-structure-recognition bottleneck that limits productivity gains. Apoidea and AWS completed the fine-tuning pipeline together, ran FinTabNet benchmarks against the base and fine-tuned versions and against Anthropic's Claude 3 Haiku and Claude 3.5 Sonnet, and published the step-by-step code on GitHub for replication. The full workload now runs on Amazon SageMaker HyperPod with vLLM hosting the quantized model. LLaMA-Factory (an open-source framework supporting over 100 LLMs with LoRA, QLoRA, SFT, RLHF and DPO) drives QLoRA fine-tuning of Qwen2-VL-7B-Instruct. SageMaker HyperPod provides the resilient training cluster powered by AWS Trainium and NVIDIA A100 / H100 GPUs, with Slurm for job scheduling, Amazon FSx for Lustre for shared storage and Amazon S3 as the data lake. Inference is served by vLLM with 4-bit quantization behind RESTful APIs. Document preprocessing converts each page to image input and HTML ground truth, and the model now preserves Chinese-language capability even though fine-tuning data is English. SuperAcc is currently deployed at more than 10 financial-services clients, supporting financial-spreading workflows where 4–6 hours of manual review per case drops to roughly 10 minutes plus less than 30 minutes of staff review. The new TEDS score of 81.1 (vs 23.4 for the base model and 69.9 for Claude 3 Haiku) lets SuperAcc handle complex multi-page financial tables at scale. Apoidea and AWS see the fine-tuning recipe as a template for other specialised information-extraction tasks in banking and finance, with the GitHub repository published to let researchers and developers replicate the work on domain-specific datasets.
Continue exploring AI deployments in the catalog.
Back to use casesCity
Hong Kong
Company/Organization
Apoidea Group
Continent
Europe
Country
Hong Kong
Category
Banks
Type
Experiment
Id
ebb16182-bbaa-4180-9e36-582efc9fc6b6
Created At
2026-08-18T21:54:29.013193+00:00