Operationalizing Model Context Protocol for Multi-Platform Academic Search: A Comparative Study of LLM Grounding

Authors

  • RIZQULLAH ARYAPUTRA PILIANG Universitas Pertahanan
  • ANINDITO Universitas Pertahanan
  • ERYAN AHMAD FIRDAUS Universitas Pertahanan

DOI:

https://doi.org/10.25134/ilkom.v20i1.513

Keywords:

Model Context Protocol (MCP), Academic paper search, Large language model integration, Retrieval-augmented generation (RAG), Preprint and metadata aggregation

Abstract

This study introduces the Paper Search Model Context Protocol (MCP) server, a unified architecture enabling Large Language Models (LLMs) to autonomously search and retrieve academic literature across heterogeneous platforms (arXiv, PubMed, bioRxiv, medRxiv, Semantic Scholar, and CrossRef). By normalizing diverse metadata into a standardized schema, the system facilitates platform-agnostic tool use. We evaluate this infrastructure by tasking GPT-4.1 and GPT-5.1 with generating grounded literature reviews on AI applications in medicine, cybersecurity, and transportation, measuring performance against retrieved abstracts using ROUGE and BERTScore. Results indicate that GPT-4.1 achieves superior semantic and structural alignment, particularly when grounded via CrossRef (BERTScore F1 = 0.881, ROUGE-L F1 = 0.375), outperforming GPT-5.1 across most metrics. While GPT-5.1 demonstrates higher unigram recall (ROUGE-1 = 0.412 on arXiv), it exhibits lower structural fidelity. These findings validate MCP as a robust integration layer for academic RAG systems and demonstrate that standardized tool interfaces enable precise, quantitative assessment of LLM grounding capabilities.

Downloads

Download data is not yet available.

References

[1] N. R. Sivakumar, “Kademlia hash snow ablation resource optimized stride scheduling for mobile computing services in healthcare sector,” Scientific Reports, vol. 15, no. 1, p. 42819, 2025.

[2] M. Elsaigh et al., “Comparative Effectiveness of Artificial Intelligence Versus Conventional Methods for Detecting Peritoneal Metastasis in Colorectal Cancer: A Systematic Review,” Cureus, 2025, doi: 10.7759/cureus.95484.

[3] T. S. T. Jakobsen, E. C. Coppulo, S. Rasmussen, and M. E. Benros, “Evaluating large language models for predicting psychiatric acute readmissions from clinical notes of population-based EHR,” medRxiv, Nov. 2025, doi: 10.1101/2025.03.07.25323558.

[4] K. Palaniappan, E. Y. T. Lin, and S. Vogel, “Global Regulatory Frameworks for the Use of Artificial Intelligence (AI) in the Healthcare Services Sector,” Healthcare, vol. 12, no. 5, p. 562, 2024.

[5] V. A. Ibiam, L. E. Omale, and O. Taiwo, “The role of Artificial Intelligence models in clinical decision support for infectious disease diagnosis and personalized treatment planning,” Int. J. Sci. Res. Anal., vol. 14, no. 3, 2025.

[6] OpenAI, “Model Context Protocol: A standard for connecting AI models to tools and data,” 2024. [Online]. Available: https://modelcontextprotocol.io

[7] A. Agarwal et al., “Tool-augmented language models in safety-critical applications: Challenges and design patterns,” arXiv preprint, arXiv:2407.12345, 2024.

[8] S. Lee, J. Park, and H. Kim, “Secure and interoperable agent communication for enterprise AI systems,” in Proc. IEEE Int. Conf. Web Services, 2024, pp. 101–110.

[9] A. S. R. Ramadhan and A. A. Fauzi, “Analisis kinerja sistem informasi akademik menggunakan pendekatan load testing dan web performance metrics,” Nuansa Informatika, vol. 16, no. 2, pp. 101–110, 2022. [Online]. Available: https://journal.fkom.uniku.ac.id/index.php/ni

[10] R. Setiawan and D. P. Sari, “Penerapan metode TOPSIS pada sistem pendukung keputusan penentuan prioritas bantuan sosial,” Nuansa Informatika, vol. 17, no. 1, pp. 45–54, 2023. [Online]. Available: https://journal.fkom.uniku.ac.id/index.php/ni

[11] M. H. Firmansyah, N. Kurniawati, and Y. A. Pratama, “Perancangan sistem informasi perpustakaan berbasis web untuk peningkatan layanan akademik,” Nuansa Informatika, vol. 15, no. 1, pp. 25–34, 2021. [Online]. Available: https://journal.fkom.uniku.ac.id/index.php/ni

[12] W. Xing et al., “MCP-Guard: A Defense Framework for Model Context Protocol Integrity in Large Language Model Applications,” arXiv preprint arXiv:2508.10991, 2025.

[13] H. Song et al., “Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem,” arXiv preprint arXiv:2506.02040, 2025.

[14] L. M. V. da Silva, A. Köcher, and F. Gehlhoff, “Beyond Formal Semantics for Capabilities and Skills: Model Context Protocol in Manufacturing,” in Proc. IEEE Int. Conf. Emerging Technol. Factory Autom. (ETFA), 2025, pp. 1–4.

[15] W. Song et al., “Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models,” arXiv preprint arXiv:2508.12566, 2025.

[16] R. V. K. Bevara et al., “Prospects of Retrieval Augmented Generation (RAG) for Academic Library Search and Retrieval,” Inf. Technol. Libr., vol. 44, no. 2, 2025.

[17] H. An, A. Narechania, E. Wall, and K. Xu, “vitaLITy 2: Reviewing Academic Literature Using Large Language Models,” arXiv preprint arXiv:2408.13450, 2024.

[18] B. Edelman and J. Skolnick, “Valsci: an open-source, self-hostable literature review utility for automated large-batch scientific claim verification using large language models,” BMC Bioinformatics, vol. 26, no. 1, 2025.

[19] N. Paulhe et al., “PeakForest: a multi-platform digital infrastructure for interoperable metabolite spectral data and metadata management,” Metabolomics, vol. 18, no. 6, p. 38, 2022.

[20] A. M. A. Zeyad and A. Biradar, “Advancements in the Efficacy of Flan-T5 for Abstractive Text Summarization: A Multi-Dataset Evaluation Using ROUGE and BERTScore,” in Proc. IEEE Asia-Pacific Conf. Comput. Sci. Data Eng. (CSDE), 2024.

[21] S. Kumar, A. Solanki, and N. Jhanjhi, “ROUGE-SS: A New ROUGE Variant for the Evaluation of Text Summarization,” Recent Adv. Comput. Sci. Commun., vol. 17, 2024.

[22] T. Zhang et al., “BERTScore: Evaluating Text Generation with BERT,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2020.

[23] N. Yadav and D. Gopinathan, “Semantic Search and Retrieval-Augmented Generation for Academic Literature Analysis Using ColBERT,” Adv. Data Sci. Adapt. Anal., 2025, doi: 10.1142/s2424922x25400017.

[24] B. Lund, “Prospects of Retrieval Augmented Generation (RAG) for Academic Library Search and Retrieval,” SSRN Electron. J., 2025, doi: 10.2139/ssrn.5295044.

[25] J. Nan, “Z-SPACE: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation,” SSRN Electron. J., 2025, doi: 10.2139/ssrn.5896270.

[26] B. Li, J. Conen, and F. Aller, “AID-Agent: An LLM-Agent for Advanced Extraction and Integration of Documents,” in Proc. 1st Workshop Res. Agent Lang. Models (REALM), 2025, pp. 80–88.

[27] W. Hu et al., “Removal of Hallucination on Hallucination: Debate-Augmented RAG,” in Proc. 63rd Annu. Meeting Assoc. Comput. Linguist. (ACL), 2025, pp. 15839–15853.

[28] A. Tay, “When is a Hallucination Not a Hallucination? The Role of Implicit Knowledge in RAG,” Rogue Scholar, 2025, doi: 10.59350/2kf7k-r1m79.

[29] Y. Zhao, Z. Liu, Y. Zheng, and K.-Y. Lam, “Attribution Techniques for Mitigating Hallucination in RAG-based Question-Answering Systems: A Survey,” TechRxiv, 2025, doi: 10.36227/techrxiv.175036754.46793269/v1.

Downloads

Published

20-01-2026

How to Cite

PILIANG, R. A., ANINDITO, & ERYAN AHMAD FIRDAUS. (2026). Operationalizing Model Context Protocol for Multi-Platform Academic Search: A Comparative Study of LLM Grounding. NUANSA INFORMATIKA, 20(1), 59–71. https://doi.org/10.25134/ilkom.v20i1.513

Similar Articles

<< < 1 2 3 4 5 6 

You may also start an advanced similarity search for this article.