Under Review
The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-context conflict can steer generative models toward confidently produced answers grounded in topically relevant but jurisdictionally incorrect, sources. Tobacco and nicotine control regulations vary across the United States, from federal, state, and municipal statutes and administrative codes that may share similar language. Thus, robust reasoning requires identifying which jurisdiction's law governs a given product, not merely retrieving semantically similar or relevant text for a given query. This is critical because emerging nicotine products (e.g., e-cigarettes and oral nicotine products) may circumvent existing regulations due to product definitions, and governments, agencies, and public health researchers and practitioners struggle to know what products are subject to which regulation.
State-of-the-art (SOTA) document retrieval-augmented generation (RAG) methods break down in this task, rife with inter-context conflict, for two reasons. First, they provide no mechanism to ground a query in a product image, and second, embedding-similarity or entity-based retrieval cannot distinguish a jurisdiction's statute from another jurisdiction's similar text. To expose this failure case and provide a benchmark for evaluating solutions to it, we introduce the Nicotine Product and Regulation Image-and-Text Surveillance Multimodal (NicoPRISM) dataset, developed by our team of computer scientists and public health researchers. NicoPRISM comprises 161,563 images from web and social media sources paired with seven structured captions, a curated knowledge base of product, health, and static legislative documents spanning 13 US exemplar jurisdictions, and 1,495 validated question-answer pairs organized into two benchmark tasks: policy compliance QA and product knowledge QA.
To address the jurisdictional disambiguation problem directly, we propose PRISM-RAG, a multimodal hypergraph RAG framework that constructs a compact hypergraph over images, captions, and entities from heterogeneous document data without any large language model calls at index time. PRISM-RAG grounds every query in a product image, then routes retrieval through a jurisdiction-aware context assembly mechanism that guarantees statutory text from the queried jurisdiction reaches the language model by construction, independent of embedding-space topology. PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p < 0.001), while requiring zero large language model calls at index time and a single call at query time. Across metrics measuring ground-truth keyword retrieval, semantic similarity, jurisdiction retrieval accuracy, compliance label accuracy, and context coverage, several of which expose limitations in SOTA document RAG methods, PRISM-RAG is competitive with and outperforms current SOTA document RAG frameworks while minimizing LLM calls at both index and query time.
PRISM-RAG addresses cross-modal grounding and jurisdictional disambiguation by constructing its retrieval structure once, offline, rather than resolving these problems at every query. A multimodal hypergraph is built over product images, structured captions, and heterogeneous product, health, and legislative documents, connecting marketing language to statutory terminology before any question is ever asked.
First, this multimodal knowledge base construction requires no large language model API calls. Entity extraction is run through a lightweight and local natural-language-processing pipeline, with no API call per document, reducing construction costs considerably. Second, product attribute nodes and legislative entity nodes are jointly clustered into concept hyperedges, the structure that bridges marketing language, e.g., “cool mint,” to statutory terminology, e.g., “characterizing flavor,” meaning this language bridging is resolved once at index time with no re-derivation required at query time. Third, with this multimodal hypergraph, query-specific sub-hypergraphs are derived from the knowledge base, and not recreated at query time, greatly reducing operations needed for response generation.
At query time, an image and a question are given as input to the PRISM-RAG pipeline. A bimodal encoder embeds both, and this joint embedding seeds retrieval over product nodes, followed by hyperedge traversal over the hypergraph built during indexing. A jurisdiction-aware context assembly step follows, injecting statutory text from the queried jurisdiction directly into the context whenever one is detected, regardless of embedding-space geometry. The generator receives this assembled context, together with the image and question, and produces the final response.
NicoPRISM is a large-scale, multimodal, expert-validated dataset for tobacco and nicotine product surveillance. It comprises web and social media images, seven-category structured attribute captions, a curated knowledge base of product, health, and legislative documents, and expert-validated question-answer pairs, organized into two benchmark tasks: policy compliance QA and product knowledge QA.
The images in this dataset were collected to support regulatory compliance surveillance and public health research. They are not intended to promote or endorse any product.
| Method | KW-F1 | CC | JA | CA | RG | Judge |
|---|---|---|---|---|---|---|
| StandardRAG | 0.182 | 0.394 | 0.450 | 0.391 | 0.734 | 2.74 |
| HippoRAG2 | 0.035 | 0.578 | 0.141 | 0.084 | 0.473 | 1.41 |
| HyperGraphRAG | 0.148 | 0.282 | 0.877 | 0.013 | 0.885 | 2.73 |
| PRISM-RAG (ours) | 0.200 | 0.550 | 0.939 | 0.418 | 0.729 | 2.86 |
All pairwise gaps against baseline methods are evaluated with bootstrap 95% confidence intervals and BH-corrected paired permutation tests (see paper for full statistical analysis).
PRISM-RAG retrieves DC-specific statutory text through its jurisdiction-aware direct injection path (JA = 1.0 for this instance). StandardRAG instead retrieves semantically similar federal content, which is grounded but answers the wrong jurisdiction's question.
This work is partly supported by the National Science Foundation (NSF) under Award No. 2501021. We also acknowledge the Arkansas High Performance Computing Center for providing computational resources.
Citation information will be added once the manuscript has been submitted.