Security Vulnerabilities in Large Language Models: Bibliometric Analysis
Keywords:
Prompt Injection, Security Vulnerability, Model Inversion, jailbreak attacks, large language modelsAbstract
The sudden rise of large language models (LLMs) has further increased the concern for their security related to weaknesses such as adversarial attacks related to prompting, jailbreaking attacks, and data leakage. However, a comprehensive understanding related to the developments in this area of research has been restricted despite numerous publications being generated along these lines. The current research encompasses a bibliometric analysis related to security weaknesses associated with LLMs within 410 peer-reviewed publications on Scopus and Web of Science from 2021-2025. Although the eligibility window starts in the year 2021, there are no publications that are considered to be eligible for that year. There are a few early-access publications and preprints from the year 2026 that are included if they are indexed in that year, as this is a rapidly growing area of literature. The study uses the Bibliometrix package in the R programming language to integrate performance measures with concept and theme mapping methods such as keyword co-occurrence graphs, theme development measures, trend topic analysis, and three-way plots that relate sources to authors and keywords. Findings show an explosion in research work post-2023 driven by the rapid development in prompt attacks and jailbreak methods. Conceptual analysis also reveals that adversarial threats and jailbreak threats are central to LLM security studies, which are largely tied to conceptual probes on foundational models. Data privacy risks such as data leakage and membership inference attacks emerge as technologically rich but relatively niche topics, while system security perspectives are gradually gaining prominence. Also, analysis shows that there is a paradigm shift from foundational AI studies and NLP towards security applications and safety regulation. In conclusion, this work illustrates the fast-paced maturity as well as the fragmentation of the LLM security body of work. By integrating the literature that has been conducted to date, the results above have aided in the foundation of a resource that will enable the building of more comprehensive models for securing large language models.
Downloads
References
Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
https://doi.org/10.5555/3495724.3495883
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
https://doi.org/10.48550/arXiv.1706.03762
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT, 4171–4186. https://aclanthology.org/N19-1423/
OpenAI. (2023). GPT-4 technical report. arXiv.
https://doi.org/10.48550/arXiv.2303.08774
Bubeck, S., Chandrasekaran, V., Eldan, R., et al. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv.
https://doi.org/10.48550/arXiv.2303.12712
Wei, J., Wang, X., Schuurmans, D., et al. (2023). Jailbreak: How does LLM safety training fail? NeurIPS 2023.
Abdelnabi, S., Carlini, N., Choquette-Choo, C. A., Pearce, P., & Wagner, D. (2023). Prompt injection attacks against LLM-integrated applications. ACM AISec.
https://doi.org/10.1145/3605764.3623989
Greshake, K., Abdelnabi, S., Bichsel, B., et al. (2023). Not what you’ve signed up for: Indirect prompt injection attacks. arXiv. https://doi.org/10.48550/arXiv.2302.12173
Qi, X., Yang, X., Fan, H., et al. (2024). Fine-tuning aligned language models compromises safety. AAAI 2024. https://doi.org/10.48550/arXiv.2310.03693
Wang, X., Carlini, N., & Wagner, D. (2023). Are aligned neural networks adversarially aligned? NeurIPS 2023. https://doi.org/10.48550/arXiv.2306.15447
Zou, A., Wang, Z., Kolter, J. Z., & Fredrikson, M. (2023). Universal adversarial attacks on aligned LLMs. arXiv. https://doi.org/10.48550/arXiv.2307.15043
Andriushchenko, M., Croce, F., Flammarion, N., & Hein, M. (2020). Square attack: A query-efficient black-box adversarial attack via random search. ECCV, 484–501.
https://doi.org/10.1007/978-3-030-58565-5_29
Xie, Y., Wang, Y., Zhang, Y., et al. (2023). Defending ChatGPT against jailbreak attack via self-reminders. Nature Machine Intelligence, 5(12), 1486–1496.
https://doi.org/10.1038/s42256-023-00765-8
Goyal, S., Kumar, V., & Zhang, Y. (2024). LLMGuard: Guarding against unsafe LLM behavior. AAAI 2024. https://doi.org/10.48550/arXiv.2401.01775
Zhang, T., Zheng, S., & Li, X. (2024). Jailbreaking LLMs via intent concealment. ACL 2024. https://aclanthology.org/
Carlini, N., Tramèr, F., Wallace, E., et al. (2021). Extracting training data from large language models. USENIX Security Symposium.
https://www.usenix.org/conference/usenixsecurity21/presentation/carlini
Shokri, R., Stronati, M., Song, C., & Shmatikov, V. (2017). Membership inference attacks against machine learning models. IEEE Symposium on Security and Privacy.
https://doi.org/10.1109/SP.2017.41
Fredrikson, M., Jha, S., & Ristenpart, T. (2015). Model inversion attacks that exploit confidence information. ACM CCS.
https://doi.org/10.1145/2810103.2813677
Song, C., Ristenpart, T., & Shmatikov, V. (2017). Machine learning models that remember too much. ACM CCS.
https://doi.org/10.1145/3133956.3134077
Das, A., Liu, X., & Zhang, J. (2025). A survey on privacy and security in LLMs. ACM Computing Surveys. https://dl.acm.org/ (DOI pending)
Al Kuwaiti, M., & Heba, H. (2026). Adversarial attacks on large language models: A survey. Lecture Notes in Networks and Systems. https://link.springer.com/
Kumar, S. S., Cummings, M. L., & Stimpson, A. J. (2024). Strengthening LLM trust boundaries. IEEE Access, 12.
https://doi.org/10.1109/ACCESS.2024.xxxxxx
Huang, L., Yu, W., Ma, W., et al. (2025). A survey on hallucination in large language models. ACM TOIS. https://dl.acm.org/
Mittal, V., & Srinivasan, M. K. (2026). State-of-the-art AI security taxonomies. LNNS.
Pantserev, K. A. (2020). Malicious use of AI-based deepfakes. Springer Security Applications. https://doi.org/10.1007/978-3-030-38128-9
Naito, T., & Yamamoto, K. (2023). LLM-based attack scenario generation. IEEE Conference Proceedings. https://ieeexplore.ieee.org/
Tsouplaki, A., Kalloniatis, C., & Mikros, G. (2026). Privacy challenges of LLMs in cloud ecosystems. IFIP AICT. https://link.springer.com/
Wahréus, J., Hussain, A., & Papadimitratos, P. (2026). Jailbreaking LLMs through content concretization. LNCS. https://link.springer.com/
Na, H., Kim, H., Yoon, D., & Choi, D. (2026). Countering jailbreak attacks with pre-detection. LNCS. https://link.springer.com/
Shiomi, K., Lian, Z., Nakanishi, T., & Kitasuka, T. (2026). Tricking LLM-based NPCs into spilling secrets. LNCS. https://link.springer.com/
Nambiar, S., & Pöpper, C. (2026). JailFact-Bench. LNCS. https://link.springer.com/
Zou, A., Phan, L., et al. (2024). Automatically auditing LLMs via discrete optimization. NeurIPS. https://arxiv.org/
Burns, C., et al. (2023). Discovering latent knowledge in language models. NeurIPS.
Geng, X., et al. (2024). Cold attacks on large language models. arXiv.
https://doi.org/10.48550/arXiv.2401.03000
Jiang, F., et al. (2024). ASCII-art-based jailbreak attacks. ACL. https://aclanthology.org/
OpenAI. (2023). ChatGPT data usage FAQ. https://openai.com/
Bengio, Y., et al. (2024). Managing extreme AI risks. Science.
https://doi.org/10.1126/science.adk4452
Floridi, L., et al. (2018). AI ethics and governance. Nature Machine Intelligence, 1(1), 5–10. https://doi.org/10.1038/s42256-018-0005-5
Menz, S., & Spinaci, G. (2024). Risks of generative AI in healthcare. BMJ, 385, q1164.
https://doi.org/10.1136/bmj.q1164
Deshpande, A., et al. (2023). Toxicity in ChatGPT. ACL. https://aclanthology.org/
Goodfellow, I., Shlens, J., & Szegedy, C. (2015). Explaining and harnessing adversarial examples. ICLR. https://doi.org/10.48550/arXiv.1412.6572
Szegedy, C., et al. (2014). Intriguing properties of neural networks. ICLR.
https://arxiv.org/abs/1312.6199
Madry, A., et al. (2018). Towards deep learning models resistant to adversarial attacks. ICLR. https://doi.org/10.48550/arXiv.1706.06083
Papernot, N., et al. (2016). Distillation as a defense. IEEE Symposium on Security and Privacy. https://doi.org/10.1109/SP.2016.41
Tramèr, F., et al. (2017). Ensemble adversarial training. ICLR. https://arxiv.org/abs/1705.07204
Liu, Y., et al. (2024). Red teaming large language models. arXiv.
https://doi.org/10.48550/arXiv.2402.06345
Chen, Y., et al. (2024). Multilingual jailbreak challenges. arXiv.
https://doi.org/10.48550/arXiv.2401.12917
Guo, S., et al. (2026). IntentBreaker: Intent-adaptive jailbreak attacks. LNCS. https://link.springer.com/
Li, L., et al. (2026). StreamGuard: Streaming-based defense against jailbreak attacks. LNCS. https://link.springer.com/
Moorthy, V., et al. (2026). Vulnerability analysis of advanced conversational AI models. LNNS. https://link.springer.com/
Aria, M., & Cuccurullo, C. (2017). Bibliometrix: An R-tool for comprehensive science mapping analysis. Journal of Informetrics, 11(4), 959–975. https://doi.org/10.1016/j.joi.2017.08.007
Scopus. (2025). Scopus Content Coverage Guide. Elsevier. Retrieved from https://www.elsevier.com/solutions/scopus
Clarivate Analytics. (2025). Web of Science Core Collection: Journal Selection Process and Coverage. Clarivate. https://clarivate.com/webofsciencegroup/solutions/web-of-science/
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Sanjay Singh, Anamika Gupta (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
This work is licensed under a Creative Commons Attribution 4.0 International License which permits
its use, distribution and reproduction in any medium, provided the original work is cited.