Many biomedical articles reference multiple datasets across different public repositories, complicating accurate metadata capture and downstream re-use. Building on our prior grounded large language model (LLM) workflows for biomedical entity annotation, we extend the approach to identify and annotate all datasets referenced by a paper, even when distributed across repositories, by combining a specialized metadata schema with a three-step, search-augmented prompting strategy.