An end-to-end evaluation system for LLM travel agents — corpus construction, question generation, answer keys grounded in mechanically decidable attributes, automated answer-quality scoring, and leaderboards — built to produce rankings that survive re-runs, addressing the non-reproducibility of LLM-as-judge grading. Reliability was validated through sample-size and ranking-stability analysis (130 questions × 5 models × 3 runs = 1,950 evaluations). First-author manuscript in preparation, with submission planned for late 2026.
A benchmark testing whether generated videos preserve evidence and maintain a consistent 3D world across camera departure, occlusion, viewpoint changes, and revisits. Controlled synthetic data is generated in Unreal Engine 5 (scripted camera trajectories, revisit paths, per-object segmentation) so that failures are attributable to the model rather than to uncontrolled scene variation. Manuscript in preparation for submission.
A stateful shopping agent for multi-turn search over a frozen 50,000-product Amazon catalog. Seekly combines reversible constraint memory, intent-aware clarification, exact and fielded BM25 retrieval, weighted reciprocal-rank fusion, and progressive exposure. Its deterministic, fully offline configuration found all 200 public targets, placing 190 at rank 1 (Hit Rate@10 1.000, MRR 0.965853, MTTC 2.305) with zero token cost. Explore the live demo or view the source on GitHub.
A RAG support-agent backend in Spring Boot with top-K retrieval and query normalisation, plus a no-evidence refusal gate that short-circuits generation when retrieval returns zero hits — trading answer coverage for groundedness so the agent declines rather than fabricates. A load-testing harness (500 requests @ concurrency 20) diagnosed a JDBC connection-pool ceiling, raising throughput 3.5 → 8.3 req/s.
A 4-stage cascaded multimodal pipeline for four-way review-quality classification on Google location reviews: rule-based filtering → FastBERT with entropy-based confidence routing → Qwen3-8B + CLIP text–image alignment → RAG verification over a FAISS index, served via vLLM. Reaches ~92% accuracy at ~150ms/sample — roughly 5 points above the latency-matched blend of the two endpoint models.
CatPals is a brownfield group project adapted from the original contact-management application AddressBook Level 3 (AB3), targeted for NUS Cate Cafe CCA volunteers to manage information about stray cats on the NUS campus.
This OpenGL portfolio is a curated collection of my CS3241 coursework, built in C++ with FreeGLUT on Visual Studio 2022. Across five assignments, I progressed from core 2D transformations to animation systems, lighting and shading, Bezier curve modeling, and finally ray tracing with shadows and texture mapping.
It showcases two major projects completed for CS3242, demonstrating practical applications of 3D modeling theory, animation principles, and computational geometry. Assignment 1 focuses on storytelling through 3D animation using Blender, while Assignment 2 implements fundamental mesh processing algorithms in C++/OpenGL.