PRISM-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Tobacco Product and Legislative Policy Reasoning

Manuel Serna-Aguilera1, Raegan Anderes2, Page Dobbs3, Khoa Luu1
1Electrical Engineering & Computer Science Department, University of Arkansas, 2Department of Health, Human Performance and Recreation, University of Arkansas, 3Fay W. Boozman College of Public Health, University of Arkansas for Medical Sciences
Under Review IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

Abstract

The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-context conflict can steer generative models toward confidently produced answers grounded in topically relevant but jurisdictionally incorrect, sources. Tobacco and nicotine control regulations vary across the United States, from federal, state, and municipal statutes and administrative codes that may share similar language. Thus, robust reasoning requires identifying which jurisdiction's law governs a given product, not merely retrieving semantically similar or relevant text for a given query. This is critical because emerging nicotine products (e.g., e-cigarettes and oral nicotine products) may circumvent existing regulations due to product definitions, and governments, agencies, and public health researchers and practitioners struggle to know what products are subject to which regulation.

State-of-the-art (SOTA) document retrieval-augmented generation (RAG) methods break down in this task, rife with inter-context conflict, for two reasons. First, they provide no mechanism to ground a query in a product image, and second, embedding-similarity or entity-based retrieval cannot distinguish a jurisdiction's statute from another jurisdiction's similar text. To expose this failure case and provide a benchmark for evaluating solutions to it, we introduce the Nicotine Product and Regulation Image-and-Text Surveillance Multimodal (NicoPRISM) dataset, developed by our team of computer scientists and public health researchers. NicoPRISM comprises 161,563 images from web and social media sources paired with seven structured captions, a curated knowledge base of product, health, and static legislative documents spanning 13 US exemplar jurisdictions, and 1,495 validated question-answer pairs organized into two benchmark tasks: policy compliance QA and product knowledge QA.

To address the jurisdictional disambiguation problem directly, we propose PRISM-RAG, a multimodal hypergraph RAG framework that constructs a compact hypergraph over images, captions, and entities from heterogeneous document data without any large language model calls at index time. PRISM-RAG grounds every query in a product image, then routes retrieval through a jurisdiction-aware context assembly mechanism that guarantees statutory text from the queried jurisdiction reaches the language model by construction, independent of embedding-space topology. PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p < 0.001), while requiring zero large language model calls at index time and a single call at query time. Across metrics measuring ground-truth keyword retrieval, semantic similarity, jurisdiction retrieval accuracy, compliance label accuracy, and context coverage, several of which expose limitations in SOTA document RAG methods, PRISM-RAG is competitive with and outperforms current SOTA document RAG frameworks while minimizing LLM calls at both index and query time.

PRISM-RAG

PRISM-RAG addresses cross-modal grounding and jurisdictional disambiguation by constructing its retrieval structure once, offline, rather than resolving these problems at every query. A multimodal hypergraph is built over product images, structured captions, and heterogeneous product, health, and legislative documents, connecting marketing language to statutory terminology before any question is ever asked.

PRISM-RAG index-time construction

First, this multimodal knowledge base construction requires no large language model API calls. Entity extraction is run through a lightweight and local natural-language-processing pipeline, with no API call per document, reducing construction costs considerably. Second, product attribute nodes and legislative entity nodes are jointly clustered into concept hyperedges, the structure that bridges marketing language, e.g., “cool mint,” to statutory terminology, e.g., “characterizing flavor,” meaning this language bridging is resolved once at index time with no re-derivation required at query time. Third, with this multimodal hypergraph, query-specific sub-hypergraphs are derived from the knowledge base, and not recreated at query time, greatly reducing operations needed for response generation.

Query-Time Inference

At query time, an image and a question are given as input to the PRISM-RAG pipeline. A bimodal encoder embeds both, and this joint embedding seeds retrieval over product nodes, followed by hyperedge traversal over the hypergraph built during indexing. A jurisdiction-aware context assembly step follows, injecting statutory text from the queried jurisdiction directly into the context whenever one is detected, regardless of embedding-space geometry. The generator receives this assembled context, together with the image and question, and produces the final response.

1
Image + Question
Product image and a natural-language query.
→
2
Bimodal Encoding
Image and text are jointly embedded.
→
3
Seed Retrieval
Top-k candidate product nodes are identified.
→
4
Hyperedge Traversal
Concept hyperedges connect attributes to legislative entities.
→
5
Jurisdiction-Aware Context Assembly
Statutory text from the queried jurisdiction is injected directly, independent of embedding similarity.
→
6
Generation
The assembled context, image, and question are passed to the generator.

NicoPRISM

NicoPRISM is a large-scale, multimodal, expert-validated dataset for tobacco and nicotine product surveillance. It comprises web and social media images, seven-category structured attribute captions, a curated knowledge base of product, health, and legislative documents, and expert-validated question-answer pairs, organized into two benchmark tasks: policy compliance QA and product knowledge QA.

161,563
Images
13
US Jurisdictions
1,495
Expert-Validated QA Pairs
7
Caption Categories
NicoPRISM dataset visualization
NicoPRISM dataset statistics

The images in this dataset were collected to support regulatory compliance surveillance and public health research. They are not intended to promote or endorse any product.

Policy Compliance QA

93.9%
of policy compliance queries are answered with passages from the correct jurisdiction — a 48.6 percentage point advantage over standard dense retrieval (p < 0.001).
Policy compliance QA results on 1,325 test instances spanning 13 jurisdictions. KW-F1 = keyword F1; CC = context coverage; JA = jurisdiction accuracy; CA = compliance accuracy; RG = response groundedness; Judge = LLM-as-judge score (1–5).
Method KW-F1 CC JA CA RG Judge
StandardRAG 0.1820.3940.4500.3910.7342.74
HippoRAG2 0.0350.5780.1410.0840.4731.41
HyperGraphRAG 0.1480.2820.8770.0130.8852.73
PRISM-RAG (ours) 0.2000.5500.9390.4180.7292.86

All pairwise gaps against baseline methods are evaluated with bootstrap 95% confidence intervals and BH-corrected paired permutation tests (see paper for full statistical analysis).

Retrieving the Right Jurisdiction's Statute

Query product image
Query: Are nicotine pouches regulated as tobacco products under Washington, D.C.'s statutes, or do they fall under a separate nicotine product framework?
PRISM-RAG KW-R: 0.40   KW-F1: 0.316   Judge: 5.0/5   JA: 1.0
Based on the DC statutes described in the provided knowledge-base excerpts, nicotine pouches are treated under the District's “tobacco products” framework, not a separate nicotine-only framework. District law's flavored-tobacco ban applies to “tobacco products,” defined to include “consumable tobacco or synthesized nicotine products” (D.C. Official Code §7-1721.01(1) / §7-1721.08). The ban also explicitly references “flavored nicotine or synthetic nicotine products” as covered flavored tobacco products (§7-1721.08(b)). So, nicotine pouches — including flavored ones — fall within DC's tobacco-products regime for purposes of the flavored-tobacco restrictions, rather than a separate nicotine-product-only framework.
StandardRAG KW-R: 0.20   KW-F1: 0.131   Judge: 2.0/5   JA: 0.0
Based on the retrieved records, nicotine pouches are treated as tobacco products under federal FDA oversight (the Family Smoking Prevention and Tobacco Control Act framework), even though they typically do not contain tobacco leaf. However, the question asks specifically about Washington, DC statutes. The provided materials do not include DC-specific statutory text, so I cannot confirm whether DC classifies nicotine pouches as tobacco products under DC law or under a separate nicotine-product scheme.

PRISM-RAG retrieves DC-specific statutory text through its jurisdiction-aware direct injection path (JA = 1.0 for this instance). StandardRAG instead retrieves semantically similar federal content, which is grounded but answers the wrong jurisdiction's question.

Acknowledgements

This work is partly supported by the National Science Foundation (NSF) under Award No. 2501021. We also acknowledge the Arkansas High Performance Computing Center for providing computational resources.

National Science Foundation Logo
Arkansas High Performance Computing Center Logo

BibTeX

Citation information will be added once the manuscript has been submitted.