Research Background
The integration of large language models (LLMs) with external knowledge sources through Retrieval-Augmented Generation (RAG) has shown promise in enhancing the accuracy and relevance of generated responses. However, the adoption of Graph-based Retrieval-Augmented Generation (GraphRAG) in enterprise settings is limited by high computational costs and latency issues. This paper addresses these challenges by proposing a more efficient and scalable framework.
In the context of the supply chain industry, the ability to quickly and accurately retrieve and synthesize information from vast amounts of unstructured data is crucial. Traditional RAG systems, while effective for straightforward fact-based queries, often struggle with complex, multi-hop reasoning tasks that are common in enterprise environments. For example, questions involving policy dependencies, multi-system workflows, or legacy code migration require connecting implicit relationships across multiple documents. In such cases, standard RAG pipelines tend to return isolated snippets without an understanding of how those pieces relate, leading to brittle or incomplete outputs.
Modern ERP systems, such as those used for finance, procurement, HR, and manufacturing, generate vast volumes of structured and unstructured data across interconnected modules. Enterprise queries often involve reasoning over configuration rules, transactional dependencies, change logs, and migration guides that are distributed across documents and systems. For instance, assessing the impact of a custom code migration in SAP’s S/4HANA may require linking legacy ABAP functions with deprecation reports, compatibility matrices, and policy guidelines. Traditional RAG systems are ill-suited for this kind of multi-hop, relational reasoning. Graph-based retrieval provides a natural fit for these scenarios, as it captures structured dependencies and enables traversal-based querying across linked entities, making GraphRAG a promising solution for ERP-related applications.
The primary shortcoming of prior approaches lies in their reliance on LLMs for constructing knowledge graphs, which incurs significant computational costs and limits scalability. Additionally, the latency introduced by graph-based retrieval hampers real-time performance, making it challenging to deploy these systems in dynamic, high-stakes enterprise environments. The proposed framework aims to overcome these limitations by introducing a dependency-based knowledge graph construction pipeline and a lightweight graph retrieval strategy, thereby reducing costs and improving scalability.
Key Findings
Dependency-Based Knowledge Graph Construction
The first key innovation of this paper is the introduction of a dependency-based knowledge graph construction pipeline that leverages industrial-grade NLP libraries to extract entities and relations from unstructured text, eliminating the need for LLMs.
The method principle behind this approach is rooted in dependency grammar, which posits that a sentence’s syntactic structure can be represented as a graph of binary head–dependent relations. For instance, in the sentence “The developer refactored the Z-report for S/4HANA,” the verb “refactored” serves as the head, while “developer,” “Z-report,” and “S/4HANA” are its dependents. By extracting such structures from unstructured text, the system can build a local knowledge graph that reflects relationships like “developer refactored Z-report” and “Z-report adapted for S/4HANA.”
The experimental setup involved evaluating the performance of the dependency-based construction approach on two SAP datasets focused on legacy code migration. The results showed that this approach achieved 94% of the performance of LLM-generated knowledge graphs (61.87% vs. 65.83%) while significantly reducing cost and improving scalability. Specifically, the dependency-based construction reduced the computational cost by 70% compared to LLM-based methods, making it a more practical solution for large-scale enterprise applications.
The key design and algorithm logic of the dependency-based construction pipeline include several stages:
- DocumentParser: Input documents are converted into a unified intermediate representation using the open-source Docling library, which retains layout, tables, and structural metadata.
- HybridChunker: Documents are split into chunks using a hierarchical chunking strategy, with a maximum size of 2048 characters and 200-character overlap, to preserve semantic cohesion.
- SentenceSegmenter: Each text chunk is segmented into individual sentences using language-specific delimiters, and sentences lacking verbs are filtered out to reduce LLM calls during downstream entity/relation extraction.
- TripleExtractor: Entities and relations are extracted using SpaCy’s dependency parser, which is built for industrial use and offers high-speed performance. The DependencyExtractor converts the parsed trees into structured knowledge triples.
- EntityRelationNormalizer: Variations of the same entity and relation are normalized and standardized to ensure compatibility with the graph database.
- RelationEntityFilter: The extracted entities and relations are post-processed to conform to a pre-defined schema, if applicable.
- GraphProducer: The generic triples are transformed into a graph format compatible with the target graph database.
- KGLoader: The graph data is loaded into the designated graph database, with different loaders implemented for various destinations.
The dependency-based construction pipeline not only reduces the computational cost but also improves the efficiency of the overall process. The use of industrial-grade NLP libraries and specialized heuristics for technical text ensures robust and accurate extraction of structured knowledge from unstructured text. This makes the framework highly adaptable for diverse text, including technical documents, policy guidelines, and legacy code.
Lightweight Graph Retrieval Strategy
The second key innovation is a lightweight graph retrieval strategy that combines hybrid query node identification with efficient one-hop traversal for high-recall, low-latency subgraph extraction.
The core design of this strategy involves a two-stage retrieval process. First, a high-recall one-hop graph traversal is conducted to identify candidate nodes. Next, a dense vector-based re-ranking step using OpenAI embeddings and cosine similarity is applied to refine the result set. This approach aligns with the classical cascaded architecture in information retrieval, where an initial recall-oriented stage is followed by a precision-oriented neural re-ranker.
The experimental evaluation demonstrated that this retrieval strategy achieved up to 15% and 4.35% improvements over traditional RAG baselines based on LLM-as-Judge and RAGAS metrics, respectively. The one-hop traversal effectively retrieved semantically related nodes while keeping the candidate set size tractable, which is crucial for scaling to large enterprise graphs. The latency of the retrieval process was also significantly reduced, with the average query time decreasing by 30% compared to traditional graph-based retrieval methods.
The key components of the efficient graph retrieval process include:
- Query Entity Identification: An optimized variant of SpaCy’s noun phrase extractor is used to pinpoint key concepts within the query, and a similarity search between the full query and node embeddings retrieves the top-𝑘 (where 𝑘 = 5) relevant nodes from the graph.
- One-Hop Traversal: A high-recall one-hop graph traversal is conducted to identify candidate nodes, ensuring that semantically related nodes are retrieved efficiently.
- Dense Vector Re-Ranking: A dense vector-based re-ranking step using OpenAI embeddings and cosine similarity refines the result set, improving the precision of the retrieved nodes.
- Subgraph Extraction: The selected subgraph, along with relevant source text chunks and extracted query entities, is passed to an LLM summarizer to generate the final, focused response.
The lightweight graph retrieval strategy not only improves the efficiency of the retrieval process but also enhances the quality of the generated responses. By combining high-recall one-hop traversal with dense vector-based re-ranking, the system can deliver coherent, multi-step responses that are both accurate and explainable. This makes the framework well-suited for complex enterprise queries that require multi-hop reasoning and structured retrieval.
Application to Legacy Code Migration
The paper is the first, to the best of the authors’ knowledge, to apply GraphRAG to a real-world legacy code migration task, demonstrating significant improvements over dense-only retrieval in both qualitative and quantitative evaluations.
The application of the proposed framework to the legacy code migration task involved assessing the impact of custom code migration in SAP’s S/4HANA. The system successfully linked legacy ABAP functions with deprecation reports, compatibility matrices, and policy guidelines, providing coherent and logically connected responses. The qualitative evaluation showed that the responses generated by the GraphRAG system were more comprehensive and accurate compared to those from traditional RAG systems. Quantitatively, the system achieved a 20% increase in the F1 score for the task, indicating a significant improvement in the quality of the generated responses.
The experimental setup for the legacy code migration task included the following steps:
- Data Collection: Two SAP datasets focused on legacy code migration were used, containing a mix of unstructured text, such as code documentation, deprecation reports, and policy guidelines.
- Knowledge Graph Construction: The dependency-based construction pipeline was used to build a knowledge graph from the unstructured text, capturing entities and relations relevant to the code migration task.
- Query Processing: Queries related to the impact of custom code migration were processed using the lightweight graph retrieval strategy, with the system retrieving both individual passages and relevant subgraphs.
- Response Generation: The selected subgraph and relevant source text chunks were passed to an LLM summarizer to generate the final, focused response.
- Evaluation Metrics: The performance of the system was evaluated using both qualitative and quantitative metrics, including the F1 score, precision, and recall.
The application of the GraphRAG framework to the legacy code migration task demonstrates its potential for improving the efficiency and accuracy of code migration and system integration processes. By leveraging the structured knowledge graph and efficient retrieval strategy, the system can provide a more comprehensive and accurate view of the migration process, enabling teams to identify and address potential issues before they become critical.
Limitations
Domain-Specific Customization
While the dependency-based construction approach is domain-agnostic, it may still require some level of customization for specific use cases, particularly in highly specialized domains.
The impact of this limitation is that the system may not perform optimally in domains with unique linguistic structures or technical jargon. For example, in the context of medical or legal documents, the system might miss important nuances or relationships. To mitigate this, the authors suggest incorporating domain-specific heuristics and training the dependency parser on in-domain data. This would allow the system to better capture the specific characteristics of the target domain, improving its overall performance.
Additionally, the authors propose developing a more flexible and adaptable framework that can easily incorporate domain-specific rules and heuristics. This could involve creating a modular architecture that allows users to plug in domain-specific modules, such as specialized NLP models or rule-based systems, to enhance the performance of the dependency-based construction pipeline. By doing so, the framework can be tailored to meet the specific needs of different industries and use cases, making it a more versatile and effective solution.
Scalability of Graph Storage and Retrieval
As document collections grow, maintaining and updating the knowledge graph becomes increasingly difficult, and the system may face scalability limitations in terms of storage and retrieval efficiency.
The impact of this limitation is that the system may struggle to handle very large graphs, leading to increased latency and higher computational costs. To address this, the authors propose using distributed graph storage solutions and optimizing the graph traversal algorithms. For example, leveraging systems like GraphScope or AliGraph, which support distributed graph analytics and GNN training, could help achieve near-linear speedups on trillion-edge workloads. Additionally, implementing incremental update mechanisms and efficient indexing strategies would further enhance the system’s scalability.
The authors also suggest exploring more advanced graph storage and retrieval techniques, such as graph partitioning and distributed processing. Graph partitioning can help distribute the graph across multiple nodes, reducing the load on any single node and improving overall performance. Distributed processing, on the other hand, can leverage the power of multiple machines to perform graph operations in parallel, further reducing latency and improving efficiency. By combining these techniques, the system can handle larger and more complex graphs, making it a more practical solution for large-scale enterprise applications.
Real-Time Performance
Despite the improvements in retrieval efficiency, the system may still face challenges in achieving real-time performance, especially in high-throughput environments.
The impact of this limitation is that the system may not be suitable for interactive use cases that require immediate responses. To mitigate this, the authors suggest further optimizing the retrieval process by pre-computing and caching frequently accessed subgraphs. Additionally, integrating more advanced graph traversal algorithms, such as Personalized PageRank (PPR), could help improve real-time performance. The authors are currently developing an optimized PPR module to better support real-time workloads.
Another approach to improving real-time performance is to implement more efficient query optimization techniques. For example, the system could use query rewriting and query planning to optimize the retrieval process, reducing the number of operations required to retrieve the relevant subgraphs. Additionally, the authors suggest exploring the use of hardware accelerators, such as GPUs and TPUs, to speed up the computation of graph operations. By leveraging these techniques, the system can achieve faster and more efficient retrieval, making it a more practical solution for real-time use cases.
Practical Implications
Enhanced Decision-Making in Supply Chain Management
The proposed GraphRAG framework can significantly enhance decision-making in supply chain management by providing more accurate and coherent responses to complex queries.
For example, in the context of inventory management, the system can help managers make informed decisions by synthesizing information from multiple sources, such as supplier contracts, demand forecasts, and historical sales data. This would enable managers to better anticipate and respond to changes in the supply chain, reducing the risk of stockouts or overstocking. The system’s ability to perform multi-hop reasoning and structured retrieval makes it particularly well-suited for scenarios that involve cross-referencing multiple documents and systems.
By leveraging the structured knowledge graph and efficient retrieval strategy, the system can provide a more comprehensive and accurate view of the supply chain, enabling managers to identify and address potential issues before they become critical. For instance, the system can help managers assess the impact of supplier delays on production schedules, or evaluate the feasibility of new sourcing strategies by analyzing historical data and current market conditions. This can lead to more informed and strategic decision-making, ultimately improving the efficiency and effectiveness of the supply chain.
Improved Code Migration and System Integration
The application of the GraphRAG framework to legacy code migration demonstrates its potential for improving the efficiency and accuracy of code migration and system integration processes.
In the context of enterprise software upgrades, the system can help developers and IT teams assess the impact of custom code migration by linking legacy functions with deprecation reports, compatibility matrices, and policy guidelines. This would provide a more comprehensive and accurate view of the migration process, enabling teams to identify and address potential issues before they become critical. The system’s ability to deliver coherent, multi-step responses makes it a valuable tool for managing complex migration projects.
By leveraging the structured knowledge graph and efficient retrieval strategy, the system can help teams navigate the complexities of code migration and system integration, reducing the risk of errors and downtime. For example, the system can help developers identify deprecated functions and recommend alternative implementations, or assess the compatibility of custom code with new system versions. This can lead to more efficient and effective code migration, ultimately improving the stability and performance of the enterprise software.
Scalable and Cost-Efficient Deployment in Enterprise Environments
The proposed framework offers a scalable and cost-efficient solution for deploying GraphRAG in enterprise environments, making it a practical choice for organizations looking to integrate proprietary data into their AI systems.
By eliminating the reliance on LLMs for knowledge graph construction and introducing a lightweight graph retrieval strategy, the system significantly reduces the computational cost and improves scalability. This makes it feasible for organizations to deploy GraphRAG systems in large-scale, dynamic environments without incurring prohibitive resource requirements. The system’s adaptability and domain-agnostic nature also make it a versatile solution for a wide range of enterprise applications, from finance and procurement to HR and manufacturing.
The framework’s cost-efficiency and scalability make it an attractive option for organizations looking to leverage the power of AI and knowledge graphs without the high computational costs and latency issues associated with traditional approaches. By providing a more efficient and scalable solution, the framework can help organizations unlock the full potential of their proprietary data, enabling them to make more informed and strategic decisions across a wide range of business functions.
Source: https://arxiv.org/abs/2507.03226