Tiny Vision–Language Models (VLMs) for Automated Archaeological Artifact Interpretation
Interpreting archaeological digital data is a labor-intensive bottleneck requiring specialized expertise. While 3D photogrammetry accelerates data collection, generating meaningful narratives about past human behavior remains challenging. This project addresses this gap by developing a specialized Tiny Vision-Language Model (TinyVLM) to automate the preliminary interpretation of archaeological artifacts. We will utilize a curated dataset of high-fidelity 3D photogrammetric models of ground stone tools (hammerstones, abraders, hematite nodules, and palettes) from the Hell Gap National Historic Landmark, dating from 12,000 to 8,000 years ago. To process this, we will translate 3D meshes into multimodal training datasets and apply advanced AI optimization, including Parameter-Efficient Fine-Tuning (LoRA), Teacher-Student Knowledge Distillation, and weight quantization, to compress the model for offline inference on edge hardware. Unlike massive, cloud-dependent AI, TinyVLMs combine computer vision with natural language reasoning on low-power, portable platforms. By enabling on-device interpretation without internet connectivity, the system ensures Indigenous data sovereignty and strengthens connections between descendant Native American communities and their material heritage. Furthermore, this project serves as a powerful catalyst for ENMU’s workforce development, directly funding and training undergraduate students across Computer Science, Electronics Engineering Technology (EET), and Anthropology. By bridging Edge-AI with heritage preservation, this project establishes a robust inter-departmental collaboration, accelerates Cultural Resource Management workflows, and positions ENMU as a highly competitive leader for future federal funding (e.g., NSF, NEH) in the emerging field of Cultural AI.