CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs
Abstract
Entity matching (EM) requires fine-grained contextual understanding and domain knowledge. Recent work shows that large language models (LLMs) can serve as strong matchers across domains, but most methods either make independent pairwise decisions or rely on manually designed composite pipelines, thus lacking flexibility in realistic multi-candidate settings. At the same time, they typically ignore inference cost at scale. We formulate LLM-based EM with candidates as a cost-aware sequential decision problem and propose CaRL-EM, a reinforcement learning controller that manages LLM operations. Given the state of an anchor record, its candidate set, and the cost, CaRL-EM adaptively chooses among different operators (Match / Compare / Select / Decide) and model capacities to maximize a quality–cost objective. The policy interacts with abstract operators, allowing the same controller to be reused with different underlying LLM backends at inference time without retraining. Experiments on 7 benchmarks show that CaRL-EM (i) learns to dynamically plan the usage of inexpensive and expensive operators based on task complexity, (ii) achieves robust zero-shot transfer across diverse datasets and domains, and (iii) consistently achieves a better quality–cost trade-off than strong LLM-based baselines and manually designed pipelines, yielding a lower inference cost at comparable or higher quality.
1 Introduction
Entity matching (EM) is a core component of entity resolution pipelines, supporting data integration, knowledge base construction, and downstream analytics in multiple domains (Shahbazi et al., 2023). Given an anchor record and a set of candidates retrieved by blocking, an EM system must identify which candidates refer to the same real-world entity (Papadakis et al., 2020; Thirumuruganathan et al., 2021; Paulsen et al., 2023).
Recent work shows that large language models (LLMs) can rival supervised deep models in pairwise matching, even under zero- or few-shot prompting (Peeters and Bizer, 2023; Li et al., 2024). However, most LLM-based EM methods focus on whether an anchor matches each candidate in isolation, ignoring mutual interactions among candidates within the same candidate set and the set-level exclusivity constraint that, in clean-clean EM, at most one candidate should be selected. In addition, EM at an industrial scale involves millions of anchors and large candidate sets, making cost a major concern (Konda et al., 2016). This becomes more significant with the application of LLMs.
To address these issues, COMEM (Wang et al., 2025) use a pipeline to allow more interaction among candidates. While these pipelines improve quality, they remain treating cost as a fixed byproduct of the architecture, applying the same sequence of operations regardless of an instance’s difficulty (Chen et al., 2023; Ong et al., 2025).
In this paper, we formulate LLM-based multi-candidate EM with blocking as a cost-aware sequential decision problem and propose CaRL-EM, as shown in Figure 1, a reinforcement learning (RL) controller that manages LLM-based EM operations. We focus on the standard clean-clean setting, in which each anchor has at most one true match in its retrieved candidate set; thus, the task studied here is to select one match or None from the candidate set. For each anchor and its candidate set, CaRL-EM maintains a compact state. We detail the state representation in §3.2. At each step, the controller selects a high level operator, Match, Compare, Select, or Decide, and a model capacity, thereby deciding whether to perform low-cost local refinement, more expensive listwise selection over the current shortlist of candidates, or terminate and output a prediction. The policy is trained with cost-aware rewards under an abstract two level cost model, more details in §3.1. CaRL-EM treats these operators as black box actions annotated with abstract cost and is separated from the underlying LLMs, so that once the policy has been trained, stronger LLMs can be plugged in at test time without retraining the controller. In summary, our work makes the following contributions:
- •
Problem formulation. We formulate LLM-based multi-candidate EM with blocking as a cost-aware sequential decision problem under RL, where a policy observes the matching state and accumulated cost information, and decides which operator to apply next. To the best of our knowledge, this is the first work to cast LLM-driven EM into a cost-aware sequential RL process.
- •
CaRL-EM. We propose CaRL-EM, a cost-aware RL controller that combines Match / Compare / Select / Decide operators and two level abstract cost models to reduce unnecessary calls. The controller is independent of specific LLMs, so backends can be swapped at test time without retraining.
- •
High efficiency and performance. On 7 benchmarks spanning products, citations, and movies, CaRL-EM outperforms the best manually designed composite pipelines, achieves higher F1 score, and reduces 78% of the cost This yields a better quality cost trade-off. Compared to domain-specific supervised EM models, CaRL-EM attains approximately 89% of their performance without any fine-tuning, while reducing 94% of the expense.
2 Related Work
2.1 Traditional and Pretrained EM
Early EM mainly used string similarity and manual rules (Papadakis et al., 2020; Barlaug and Gulla, 2021). While neural models such as DeepER and DeepMatcher improved semantic capture, they remain heavily dependent on labeled data (Ebraheem et al., 2017; Mudgal et al., 2018). Recent pretrained language models further leverage cross-encoders like Ditto (Li et al., 2020) or dual-encoders (Shah et al., 2018; Tracz et al., 2020). However, these methods typically treat pairs independently, ignoring mutual interactions and global consistency. Moreover, their reliance on fine-tuning limits their transferability to new domains (Li et al., 2020; Peeters and Bizer, 2021).
2.2 LLM-Based Strategies
LLMs can perform pairwise matching in a zero- or few-shot manner with performance often rivaling supervised models, while providing explanations of their decisions Peeters et al. (2025). Most methods prompt the LLM to output a Yes/No decision for each pair, which keeps the pairwise limitation.
To go beyond independent decisions, COMEM analyzes three LLM-driven strategies for multi-candidate EM: Match, Compare, and Select Wang et al. (2025). It manually builds a pipeline that uses cheap steps to re-rank candidates and then calls a stronger LLM to select the final match. However, the pipeline structure is the same for all anchors, and the inference cost is implicitly determined by the chosen pipeline.
2.3 Cost-Aware Control with RL
LLMs often work well without much tuning, but they can be costly to run at scale. This has led to growing interest in controlling LLM inference under resource constraints. Recent work uses lightweight policies or RL to decide when to call external tools or how much context to use, trading off accuracy for token cost (Feng et al., 2025; Zhang et al., 2024). These approaches are mainly studied in question answering or reasoning tasks rather than EM. In contrast, EM tasks that require large-scale LLM calls typically rely on fixed pipelines.
We cast multi candidate EM as a cost-aware sequential decision problem. CaRL-EM learns to chooses Match, Compare, Select, and Decide, as well as LLMs with different costs, to navigate the accuracy and cost trade-off.
3 CaRL-EM: A Cost-Aware RL Controller for LLM-Based EM
We consider the standard blocked EM setting. For each anchor record , a blocking retrieves a candidate set . The goal is to decide which candidate matches , or output None. We focus on the common clean-clean scenario in which at most one candidate in is a true match (Gemmell et al., 2011). We cast multi-candidate EM for each anchor as a cost-aware sequential decision problem and learn a policy that chooses which action to apply next, balancing both expected matching quality and inference cost. All hyperparameters introduced in this section, including , , , , , and top-, are specified with their values in Appendix A.
3.1 LLM-Based Operators and Abstract Cost
CaRL-EM uses four high-level actions: three LLM-based operators (Match, Compare, Select) and a terminal action (Decide), as shown in Figure 2. The LLM-based operators are implemented with dedicated prompting templates in Appendix B.
Match. Given an anchor candidate pair , a Match call asks a lightweight matcher, Flan-T5-xl (Raffel et al., 2020) is used for Match and Compare in our experiment, to output YES/NO. We use the model probability of emitting YES to update the internal confidence score.
Compare. A Compare call presents with two candidates and asks which candidate is more likely to match ; the resulting preference is translated into a small update of their respective scores and , sharpening the score distribution.
Select. A Select call presents and a shortlist of top- candidates ranked by the current scores. It prompts the LLM to perform listwise reasoning to identify the most likely match within this group or output None. The selected candidate receives a score boost, while others in the shortlist are penalized. This listwise feedback refines the global ranking without making a final decision.
Decide. Decide is the only terminal operator. The action set includes . The policy directly chooses a candidate or None without a specific thresholding.
Operators vary significantly in complexity and input scope, often necessitating different model capacities. Local operations like Match and Compare are computationally lighter and can be handled by smaller models, whereas the listwise Select operator requires reasoning over a larger context of multiple candidates, often demanding a larger, more capable model Wang et al. (2025). This disparity in both model size and prompt length leads to distinct inference costs.
We capture this difference using a two level abstract cost model. Each operator is assigned a cost label in , which we normalize to numerical values, e.g., for low-cost local Match/Compare calls and for high-cost shortlist Select calls. We set by default. Increasing it to yields nearly the same , but increases Match/Compare usage and leads to higher cross-benchmark variance; more details are provided in Appendix C.
These numbers are not tied to any particular pricing scheme and are intended to reflect relative expense in terms of tokens and API calls. Thus, the abstract cost is a control signal for policy learning rather than an exact accounting model of any single deployment environment. This abstraction decouples the learned policy from a specific model. By training on relative cost tiers, CaRL-EM allows users to swap different underlying LLMs for the operator at inference time without retraining the controller. Given a sequence of actions for an anchor, the total abstract cost is
and this cost enters the RL objective as a penalty, encouraging the policy to use cheap operators whenever they suffice.
3.2 MDP Formulation
Actions change the candidate state and incur cost, so decisions are sequential, including when to stop with Decide. We therefore model the decision process for each anchor as an episodic Markov decision process (MDP) (Sutton et al., 1998). At each step, an action invokes an operator that updates the state, and we learn a policy to maximize the expected return until Decide or .
State. At step , the policy observes a state vector that captures the current belief state, resource consumption, and interaction history. Formally, let be the maximum number of candidates and be the length of the action history window. The state vector is constructed as the concatenation of several feature groups:
contains the current confidence scores, indicates whether each candidate is still eligible to be processed by a Match operator (e.g., has not exceeded a per-candidate call limit), and tracks the frequency of each candidate in Compare calls. The global context consists of the accumulated abstract cost, the normalized time step , and the current maximum candidate score. To detect loops or repetitive patterns, encodes the last high level actions as a one-hot vector. Optionally, dense semantic embeddings of the anchor and the current top-scoring candidate () are appended to provide grounding. During training, we add small Gaussian noise into to enhance policy robustness.
Action. At step , the controller chooses a discrete action that specifies an operator and its operands. Actions include for inspecting candidate , for evaluating a selected champion–challenger pair, for listwise selection over the current top- candidates, and terminal / . The episode ends when Decide is chosen or when is reached. Executing invokes the corresponding LLM operator, updates candidate scores and usage statistics, and accumulates abstract cost. The full equations of candidate score updating are provided in Appendix D.
Reward. We design the reward to favor correct final decisions while keeping the cost low. Let be the score of candidate at step , its label, and the active candidates. We define a global margin as the score gap between the best true match and the best non-match; if no match exists, it is the negative best score:
At a terminal step , the agent outputs and receives a correctness reward:
The total terminal reward is formulated as:
where rewards stopping before the deadline , reflects the abstract cost, and penalizes premature decisions with low confidence.
For non-terminal steps , we use shaping to guide the agent (Ng et al., 1999; Wiewiora, 2003):
3.3 Policy, Training, and Inference
CaRL-EM uses a lightweight policy network , implemented as a 3-layer MLP (768–384–192), which maps the state vector to a distribution over high level actions. We train with Proximal Policy Optimization (PPO) and a learned value baseline (Schulman et al., 2017).
At test time, the controller starts from the initial state and iteratively selects actions until it chooses Decide or reaches the step limit . Because the controller interacts only with abstract operators and their cost labels, and is decoupled from the underlying LLMs that implement them, stronger LLMs can be plugged in at inference time without retraining the policy, as long as the operator interfaces and relative cost levels are preserved. This allows CaRL-EM to benefit from future improvements in LLM capabilities while retaining the learned decision strategy.
4 Experiments
LLMs used in our experiments.
We evaluate CaRL-EM with a diverse set of LLM backends, covering both proprietary commercial APIs (OpenAI, 2024) and open weight models (Chung et al., 2024; Grattafiori et al., 2024; Agarwal et al., 2025; ERNIE Team, 2025; Team et al., 2025; Yang et al., 2025); more detail are in Table 1. For open weight models, we report a price range collected from several major inference providers11 1 Prices are taken from providers: DeepInfra https://deepinfra.com/; Fireworks AI https://fireworks.ai/; Together.ai https://www.together.ai/; Novita https://novita.ai/; Groq https://groq.com/.
| Model Name | Size | CoT | Cost ($/1M tokens) | |
| Input | Output | |||
| Flan-T5-xl | 3B | – | 0.10 | 0.10 |
| Llama-3.1 | 8B | – | ||
| GPT-oss | 20B | ✓ | ||
| ERNIE 4.5 | 21B | ✓ | 0.07 | 0.28 |
| Gemma 3 | 27B | – | 0.1 | 0.2 |
| Qwen3 | 30B | ✓ | ||
| GPT-4o mini | – | – | 0.15 | 0.60 |
Datasets and Zero-shot Transfer Setting. We train our approach on Abt-buy (AB) and evaluate it on standard entity resolution benchmarks extensively used in prior literature (Mudgal et al., 2018). These datasets cover diverse domains, including e-commerce, academic citations, movies, and restaurants. Following the experimental protocol established in COMEM (Wang et al., 2025), we employ a blocking stage using Sparkly to retrieve the top-10 most likely candidates from the target table for each anchor record. Blocking recall@10 is high across all datasets, ranging from 94.89%-to 99.96%, so remaining errors mainly reflect the quality of decisions. The final evaluation set contains 400 anchors per dataset, including 300 anchors with one true match, and the rest have no true match. This setup shifts the task from isolated pairwise classification to a realistic multi-candidate selection problem in the clean-clean setting, where the model must choose one candidate or reject all candidates.
To evaluate the generalization capability of CaRL-EM, we use a zero-shot transfer protocol. CaRL-EM is trained only once on the AB dataset and evaluates it on the other seven benchmarks without any fine-tuning or adaptation. Since AB comes from a different domain and source than the test datasets, the controller cannot learn target-domain knowledge during training, and there is no data leakage. In contrast, for supervised baselines, like Ditto (Li et al., 2020), we sample an additional 5,000 labeled pairs from each target dataset to train domain-specific models. The specific datasets are: Abt-buy (AB), Amazon-Google (AG), DBLP-ACM (DA), DBLP-Scholar (DS), IMDb-TMDb (IM), IMDb-TVDb (IT), TMDb-TVDb (TT), and Walmart-Amazon (WA).
Hardware platforms and their price. All experiments involving local computation are conducted on NVIDIA H100 GPUs. To ensure a realistic and fair economic comparison, we standardize the GPU compute cost at $5.98 per hour. This rate is derived by averaging the on-demand pricing for H100 instances across four major cloud service providers; more details are provided in Appendix E.
Evaluation Metrics. Traditional pairwise F1 scores compute performance over isolated pairs, ignoring the mutual exclusivity often required in real-world applications, such as an anchor has at most one valid match (Christen, 2012). Given that our system makes a holistic decision over the candidate set , we employ instance level metrics rather than pair level ones. This protocol is strictly more demanding: a “Success” is counted only when the exact ground-truth is selected from candidates; selecting a wrong candidate counts as both a False Positive and a False Negative. Notably, under the assumption of single-match validity, this instance level protocol aligns mathematically with the standard pairwise F1 score. Accordingly, we report to measure selection accuracy on anchors with valid matches, and to evaluate rejection sensitivity on those without. We define as their unweighted average to ensure a balanced assessment that penalizes both selecting when none exists and misses.
Cost Calculation. We perform a comprehensive economic analysis covering both computational and API costs. For LLM-based components, costs are derived from token usage based on the standardized API pricing across multiple platforms, as shown in Table 1. For the training of CaRL-EM and the full lifecycle of the supervised baseline Ditto, we incorporate GPU compute costs standardized at $5.98/hour.
All compared methods use the same candidate sets produced by the same blocking stage. Therefore, blocking does not affect the relative comparison among decision policies; in our setting, its cost is also negligible compared with LLM-based decision cost. For completeness, we report the blocking cost separately in Appendix F.
During the inference phase, CaRL-EM’s policy network is implemented as a lightweight MLP. Its computational overhead on a commodity CPU is negligible compared to the network latency of LLM API calls. Therefore, we assume zero hardware cost for CaRL-EM’s inference, as it can be efficiently served without a GPU. For CaRL-EM, we distinguish deployment-time inference cost from one-time offline training cost. The latter is incurred once; we report it transparently in Appendix F.
5 Results
5.1 Main Results
Table 2 reports instance level and cost on 7 datasets. Ditto is trained in-domain, so it is trained separately on each target dataset. In contrast, CaRL-EM is trained only once on the AB dataset and then used zero-shot on the other datasets without any additional fine-tuning. Overall, CaRL-EM keeps competitive while using less inference cost, which leads to a higher efficiency score under our metric.
| Method | Training | AG | DA | DS | IM | IT | TT | WA | Avg. | Cost ($) | Eff. Score |
| Reference: Supervised (In-domain) | |||||||||||
| Ditto (Li et al., 2020) | per-dataset | 63.3 | 96.8 | 88.4 | 93.9 | 89.8 | 87.0 | 79.8 | 85.64 | 2.22 | 73.2 |
| Zero-shot / Transfer Settings | |||||||||||
| Matching (Peeters et al., 2025) | none | 29.18 | 70.88 | 71.27 | 43.91 | 34.15 | 35.44 | 26.09 | 46.66 | 2.19 | 40.2 |
| Comparing | none | 47.35 | 85.04 | 77.89 | 61.66 | 33.60 | 46.29 | 43.71 | 58.08 | 0.54 | 134.4 |
| Selecting | none | 42.55 | 66.90 | 59.75 | 92.20 | 84.00 | 84.25 | 63.70 | 69.94 | 0.15 | 499.7 |
| COMEM (Wang et al., 2025) | none | 60.01 | 56.94 | 69.54 | 87.32 | 75.46 | 84.54 | 78.76 | 74.25 | 0.59 | 160.0 |
| CaRL-EM (Ours) | AB-only | 54.65 | 78.50 | 77.75 | 88.70 | 82.65 | 85.80 | 64.60 | 76.09 | 0.13 | 623.7 |
Comparison with the supervised baseline. Ditto is an in-domain supervised model, so it is trained separately on each dataset. CaRL-EM is trained once on AB and then transferred to the other 7 datasets with no extra tuning. Under this setting, CaRL-EM achieves an average of 76.09, while Ditto reaches 85.64. This means CaRL-EM scores about 89% of Ditto. However, the cost of CaRL-EM is 5.9% of Ditto’s cost. This shows that the CaRL-EM can reuse one policy across datasets and keep the cost low. As we discuss in §5.3, Ditto does not transfer well across domains without retraining.
Comparison with LLM-based methods. Among zero-shot methods, CaRL-EM achieves the best average . It also outperforms the strongest hand-crafted baseline, COMEM. At the same time, it uses less cost, with about a 4.5 reduction. Intuitively, a fixed pipeline is more likely to pick a candidate even when the evidence is weak, which leads to false positives. In contrast, CaRL-EM first applies the global Select operator and then uses Match and Compare for local checks. When the information is not enough, it is also more likely to stop with a no-match decision. We further support this point with in Table 3.
5.2 Cost-Efficiency Analysis
Cost quality trade-off. Figure 3 compares methods in the dataset level total cost and .22 2 We report the cost of inference for all zero-shot methods in this plot. CaRL-EM’s one-time offline cost is reported separately in Appendix F. In the figure, CaRL-EM points are mostly in the high-quality, low-cost region. Without retraining the controller, swapping in a stronger backend LLM usually improves , but it also increases cost. Under the same LLM, CaRL-EM (GPT-4o mini) achieves a higher than the state-of-the-art cost-efficient baseline COMEM, while using only 22% of its average total cost. Overall, CaRL-EM forms a new Pareto frontier (Deb et al., 2002). This shows that it offers the best quality cost trade-off among the methods we evaluate. The gain mainly comes from the controller’s adaptive decisions, rather than any single LLM.
5.3 Transfer and Robustness to LLM Backends
Zero-shot transfer. Table 3 compares CaRL-EM with COMEM and a transfer baseline based on Ditto. Ditto is trained on AB and then applied to the other datasets without retraining, so its drops from 85.64 to 59.75. In contrast, CaRL-EM is also not retrained, yet it still achieves a high . This suggests that the learned decision strategy captures potential patterns in multi-candidate EM that transfer across datasets.
Swapping LLM backends. The same controller can also work with different LLM backends. We test this by swapping the backend models at inference time, without retraining the controller. As shown in Table 3, stronger backends usually improve , but they also increase cost. The controller mainly adjusts how often it uses Match and Compare to balance quality and cost. This pattern matches our design goal: the policy learns the decision sequence, rather than overfitting to one dataset or one specific LLM.
| Backend | Avg. | Avg. | Avg. | Cost | |||
|---|---|---|---|---|---|---|---|
| Ditto† | 73.52 | 45.99 | 59.75 | – | – | – | 0.05 |
| COMEM | 85.66 | 60.78 | 73.22 | 10.00 | – | 1.00 | 0.59 |
| CaRL-EML (Llama-3.1) | 81.23 | 57.49 | 69.36 | 1.42 | 2.03 | 0.99 | |
| CaRL-EME (Gemma3) | 82.86 | 60.56 | 71.70 | 1.69 | 2.02 | 0.99 | 0.11 |
| CaRL-EM (GPT-4o mini)‡ | 83.99 | 68.20 | 76.09 | 1.92 | 1.95 | 0.99 | 0.13 |
| CaRL-EMG (GPT-oss) | 85.41 | 72.91 | 79.14 | 2.03 | 1.91 | 0.99 | |
| CaRL-EME (ERNIE 4.5) | 83.77 | 72.17 | 77.97 | 2.47 | 1.89 | 0.99 | 0.49 |
| CaRL-EMQ (Qwen3) | 85.41 | 73.40 | 79.37 | 2.18 | 1.90 | 0.99 |
5.4 Controller Behavior and Position Bias
We inspect the learned policy to understand how it manages LLM operators and how this affects robustness to position bias in long candidate lists (Liu et al., 2024).
Operator usage patterns. Figure 4 shows the average number of operator calls per anchor. Compared with COMEM, CaRL-EM uses fewer Match and Compare calls. At the same time, the amount of tool usage varies across domains and datasets, showing that the learned policy does not follow a fixed call pattern. We also observe that on relatively harder datasets such as AG and WA, CaRL-EM tends to run more steps before making a final decision. In contrast, COMEM follows a fixed pipeline and cannot adjust its strategy across instances or datasets.
Mitigating long-context position bias. We run a controlled test to isolate position effects in the 7 test datasets. For each anchor, we force the gold candidate to appear at a fixed position (0–9) in the initial list. We keep the gold position fixed and shuffle the other candidates. We repeat this process 10 times for each position and report the average result.
As shown in Figure 5, the Selecting baseline is more sensitive to where the gold candidate appears in the list. It performs notably worse when the gold candidate is placed at position 1, and it still varies across other positions. The Matching and Matching baselines are flatter, but Comparing declines slightly, while Matching increases slightly. The shaded bands indicate gaps across datasets. This gap is larger for Matching and Comparing, which suggests weaker stability across domains. In contrast, CaRL-EM stays more stable across positions and shows the smallest cross-dataset variation. A key reason is that CaRL-EM does not rely on a single long listwise call. It can flexibly combine Match/Compare and Select, and it can perform mutual verification between various operators, which reduces dependence on the initial order.
5.5 Qualitative Analysis: Policy Behavior
To understand how CaRL-EM achieves efficiency, we visualize the decision trajectory on two cases from the AG dataset, as shown in Figure 6.
Case 1 (easy): From the initial confidence scores, the gold candidate (c2) has a clear advantage. The policy first calls Select on the top-4 list. It then uses two Compare calls as quick checks. It finally decides on c2. This takes only a few steps and avoids checking all candidates one by one; in this case, the policy does not call the pairwise Match operator.
Case 2 (Hard): In this case, several candidates start with similar scores, and the gold candidate (c0) doesn’t have the highest confidence score. The policy again calls Select on the top-4 list, but the first result is c4. It then calls Match for local verification and finds that it conflicts with the result from the previous step. Therefore, it runs multiple Match calls and gradually improves the confidence of c0. This shows that the policy spends more local checks on hard cases, but stops early on easy ones, which allocates the budget based on instance difficulty. This adaptive behavior demonstrates that CaRL-EM acts as a "System 2" thinker, allocating computational resources dynamically based on instance difficulty (Kahneman, 2011).
5.6 Ablation and Design Choices
We run ablations to understand which components matter. We retrain each variant on AB and report results averaged over the other seven datasets. Table 4 summarizes accuracy, stability (Std.), and normalized inference cost.
Cost-aware reward. Removing the cost term keeps almost unchanged, but the cost increases by 15%. This shows that the cost term is needed to learn a more efficient policy.
Potential-based shaping. Removing the global potential term reduces the and value and makes performance unstable. This shows that the shaping signal helps the policy separate true and false candidates during the episode.
Operators. Each operator plays a different role. Without Match, the model is cheaper but its accuracy drops, especially on none cases. Without Select, the model is the cheapest and is high, but drops and overall performance is lower. Without Compare, performance also drops, which suggests that pairwise checks help when top candidates are close.
| Variant | Std. | Cost | |||
|---|---|---|---|---|---|
| Full (ours) | 83.99 | 68.20 | 76.09 | 12.24 | 0.13 |
| w/o cost | 82.96 | 69.16 | 76.07 | 13.05 | 0.15 |
| w/o | 83.03 | 68.89 | 75.93 | 14.56 | 0.13 |
| w/o Match | 81.91 | 49.86 | 65.87 | 12.59 | 0.10 |
| w/o Compare | 80.04 | 67.31 | 73.69 | 15.04 | 0.13 |
| w/o Select | 75.13 | 71.36 | 73.26 | 20.60 | 0.08 |
6 Conclusion
We study LLM-based EM in the practical blocked setting, where each anchor comes with a small candidate set and cost quickly becomes the bottleneck at scale. We propose CaRL-EM, a reinforcement learning controller that treats EM as a sequential decision process and decides when to use Match, Compare, Select, or Decide, while also choosing between cheaper and stronger model capacities. To our knowledge, this is the first work that casts LLM-driven EM as a cost-aware sequential RL problem.
Across seven benchmarks under zero-shot transfer, CaRL-EM learns to spend less on easy cases and check more on hard ones, and it yields a better quality cost trade-off than strong LLM baselines and hand-designed pipelines. More broadly, our formulation offers a way to think about trading off decision strategy performance and cost through a learned controller, rather than a fixed pipeline.
Limitations
Our study focuses on the clean-clean setting where an anchor has at most one true match. While this assumption fits many standard benchmarks and blocking-based pipelines, it does not cover settings with multiple valid matches, one-to-many links, or noisy/duplicate-heavy tables. Extending CaRL-EM to those cases would likely require changes to both the state (e.g., tracking multiple plausible matches) and the stopping/decision rule.
In addition, we mainly consider small candidate pools (top-10 in our protocol). When candidate sets become much larger, a controller that directly reasons over the whole list may become less effective, and a different design (e.g., multi-stage pruning, hierarchical control, or tighter coupling with retrieval) may be needed to keep both cost and decision quality under control.
Finally, CaRL-EM is trained with a coarse two level abstract cost model. This makes backend swapping simple and keeps the policy less tied to a specific pricing scheme, but it does not capture the full complexity of real deployment costs, such as prompt-length differences, token-based billing, latency constraints, batching effects, and provider-specific pricing. As a result, the learned behavior may not be cost-optimal under a different cost surface, and deployments may require recalibrating the cost function (or retraining) to match the target environment.
Ethical Considerations
We do not collect new data. Our experiments use public entity-matching benchmarks and we do not release any additional data. The main risk is incorrect matches, which can lead to wrong merges and downstream errors. High-stakes use should include auditing and human review. We discourage using entity matching to link personal identities across datasets and emphasize legal and ethical compliance.
Acknowledgments
We thank Frank van Harmelen for his valuable feedback and suggestions on this paper. We also thank the anonymous reviewers for their constructive comments. Chaohui Guo is supported by the China Scholarship Council (CSC).
References
- Gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925. Cited by: §4.
- Neural networks for entity matching: a survey. ACM Trans. Knowl. Discov. Data 15 (3). External Links: ISSN 1556-4681, Link, Document Cited by: §2.1.
- Frugalgpt: how to use large language models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176. Cited by: §1.
- Data matching: concepts and techniques for record linkage, entity resolution, and duplicate detection.. Springer: Data-centric systems and applications. Cited by: §4.
- Scaling instruction-finetuned language models. Journal of Machine Learning Research 25 (70), pp. 1–53. Cited by: §4.
- A fast and elitist multiobjective genetic algorithm: nsga-ii. IEEE Transactions on Evolutionary Computation 6 (2), pp. 182–197. External Links: Document Cited by: §5.2.
- DeepER–deep entity resolution. arXiv preprint arXiv:1710.00597. Cited by: §2.1.
- ERNIE 4.5 technical report. Note: https://yiyan.baidu.com/blog/publication/ERNIE_Technical_Report.pdf Cited by: §4.
- Retool: reinforcement learning for strategic tool use in llms. arXiv preprint arXiv:2504.11536. Cited by: §2.3.
- Improving entity resolution with global constraints. arXiv preprint arXiv:1108.6016. Cited by: §3.
- The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §4.
- Thinking, fast and slow. Farrar, Straus and Giroux, New York. External Links: ISBN 9780374275631 Cited by: §5.5.
- Magellan: toward building entity matching management systems. Proc. VLDB Endow. 9 (12), pp. 1197–1208. External Links: ISSN 2150-8097, Link, Document Cited by: §1.
- On leveraging large language models for enhancing entity resolution: a cost-efficient approach. arXiv preprint arXiv:2401.03426. Cited by: §1.
- Deep entity matching with pre-trained language models. arXiv preprint arXiv:2004.00584. Cited by: §2.1, §4, Table 2.
- Lost in the middle: how language models use long contexts. Transactions of the Association for Computational Linguistics 12, pp. 157–173. External Links: Link, Document Cited by: §5.4.
- Deep learning for entity matching: a design space exploration. In Proceedings of the 2018 international conference on management of data, pp. 19–34. Cited by: §2.1, §4.
- Policy invariance under reward transformations: theory and application to reward shaping. In Proceedings of the Sixteenth International Conference on Machine Learning, ICML ’99, San Francisco, CA, USA, pp. 278–287. External Links: ISBN 1558606122 Cited by: §3.2.
- Routellm: learning to route llms with preference data, 2024. URL https://arxiv. org/abs/2406.18665 4. Cited by: §1.
- GPT-4o mini: advancing costefficient intelligence. Note: https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Cited by: §4.
- Blocking and filtering techniques for entity resolution: a survey. ACM Comput. Surv. 53 (2). External Links: ISSN 0360-0300, Link, Document Cited by: §1, §2.1.
- Sparkly: a simple yet surprisingly strong tf/idf blocker for entity matching. Proc. VLDB Endow. 16 (6), pp. 1507–1519. External Links: ISSN 2150-8097, Link, Document Cited by: §1.
- Dual-objective fine-tuning of bert for entity matching. Proc. VLDB Endow. 14 (10), pp. 1913–1921. External Links: ISSN 2150-8097, Link, Document Cited by: §2.1.
- Using chatgpt for entity matching. In European Conference on Advances in Databases and Information Systems, pp. 221–230. Cited by: §1.
- Entity matching using large language models. In Proceedings 28th International Conference on Extending Database Technology, EDBT 2025, Barcelona, Spain, March 25-28, 2025, A. Simitsis, B. Kemme, A. Queralt, O. Romero, and P. Jovanovic (Eds.), pp. 529–541. External Links: Link, Document Cited by: §2.2, Table 2.
- Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21 (1). External Links: ISSN 1532-4435 Cited by: §3.1.
- Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. Cited by: §3.3.
- Neural network based extreme classification and similarity models for product matching. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 3 (Industry Papers), S. Bangalore, J. Chu-Carroll, and Y. Li (Eds.), New Orleans - Louisiana, pp. 8–15. External Links: Link, Document Cited by: §2.1.
- Through the fairness lens: experimental analysis and evaluation of entity matching. Proc. VLDB Endow. 16 (11), pp. 3279–3292. External Links: ISSN 2150-8097, Link, Document Cited by: §1.
- Reinforcement learning: an introduction. Vol. 1, MIT press Cambridge. Cited by: §3.2.
- Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Cited by: §4.
- Deep learning for blocking in entity matching: a design space exploration. Proc. VLDB Endow. 14 (11), pp. 2459–2472. External Links: ISSN 2150-8097, Link, Document Cited by: §1.
- BERT-based similarity learning for product matching. In Proceedings of Workshop on Natural Language Processing in E-Commerce, H. Zhao, P. Sondhi, N. Bach, S. Hewavitharana, Y. He, L. Si, and H. Ji (Eds.), Barcelona, Spain, pp. 66–75. External Links: Link Cited by: §2.1.
- Match, compare, or select? an investigation of large language models for entity matching. In Proceedings of the 31st International Conference on Computational Linguistics, pp. 96–109. Cited by: §1, §2.2, §3.1, §4, Table 2.
- Potential-based shaping and q-value initialization are equivalent. J. Artif. Int. Res. 19 (1), pp. 205–208. External Links: ISSN 1076-9757 Cited by: §3.2.
- Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: §4.
- TREACLE: thrifty reasoning via context-aware llm and prompt selection. CoRR. Cited by: §2.3.
Appendix A Hyperparameter Values
Table 5 lists the concrete hyperparameter values used in our experiments.
| Symbol | Meaning | Value |
|---|---|---|
| max #candidates | ||
| action history length | ||
| max decision steps per anchor | ||
| cost ratio | ||
| top- shortlist size(s) for Select | ||
| per-step cost penalty weight in reward | ||
| global potential shaping weight | ||
| local shaping weight for Select | ||
| shaping discount factor | ||
| Gaussian noise |
| Avg. | Avg. | Avg. | Std. | Cost | ||||
|---|---|---|---|---|---|---|---|---|
| 2.5 | 83.99 | 68.20 | 76.09 | 1.92 | 1.95 | 0.99 | 12 | 0.13 |
| 5 | 83.17 | 71.23 | 77.20 | 2.42 | 1.58 | 0.99 | 15 | 0.13 |
| Method | Blocking ($) | Training ($) | Test ($) | Total E2E ($) | Note |
|---|---|---|---|---|---|
| Ditto | 0.00503 | 2.175258222 | 0.04777355556 | 2.228061778 | per-target training |
| Matching | 0.00503 | 0 | 2.19 | 2.19503 | - |
| Comparing | 0.00503 | 0 | 0.54 | 0.54503 | - |
| Selecting | 0.00503 | 0 | 0.15 | 0.15503 | - |
| COMEM | 0.00503 | 0 | 0.59 | 0.59503 | - |
| CaRL-EM | 0.00503 | 16.45 | 0.13 | 16.58503 | train once on AB, reused across targets |
Appendix B Prompt Templates
B.1 Match
Output constraint: the first line must be exactly one token from [YES] or [NO].
Decide if the two records refer to the SAME real-world entity. Output EXACTLY ONE token on the first line: [YES] for same, [NO] for different. [ANCHOR] {anchor} [CANDIDATE] {candidate} Answer:
B.2 Compare
Output constraint: the first line must be exactly one token from [0] or [1].
Which candidate better matches the anchor? Return EXACTLY ONE token: [0] for the first candidate, [1] for the second. [ANCHOR] {anchor} [0] {a} [1] {b} Answer:
B.3 Select
Output constraint: the first line must be exactly one token from [0..n-1] or [NONE].
Return only one of [0..n-1] or [NONE]. Select the single best-matching candidate for the anchor. Answer with EXACTLY one token on the first line: [0..{n-1}] or [NONE]. Do NOT include any other words or punctuation. [ANCHOR] {anchor} OPTIONS: [0] {c0} [1] {c1} ... [n-1] {cn-1}
Appendix C Sensitivity to the Abstract Cost Ratio
We test the robustness of the two-level abstract cost ratio by changing it from to , retraining CaRL-EM on Abt-Buy (AB) and evaluating zero-shot on the seven target benchmarks. As shown in Table 6, remains nearly unchanged. However, the learned policy makes slightly more Match/Compare calls and exhibits slightly higher variance across benchmarks.
Appendix D Confidence Update Rules
We maintain per-candidate internal confidence scores .
Match update.
Let be the matcher probability of [YES]. We map it to and update candidate by exponential smoothing:
| (1) |
where if , otherwise .
Compare update.
Given a compared pair , the comparator returns a winner and a confidence . Let and set if , else . Then:
| (2) |
Select update.
Let be the current top- shortlist. If the selector outputs a valid index :
| (3) |
If the selector output is invalid (e.g., parsing failure), we penalize the shortlist:
| (4) |
Finally, all scores are clipped to .
Appendix E H100 pricing from different cloud service platforms
Table 8 lists the H100 GPU server quotes we collected from 4 common cloud service providers. According to this, we estimated the average price of H100 servers on the market.
| Provider | Instance | GPUs per instance | USD / GPU*h |
|---|---|---|---|
| AWS | p.548xlarge | 8 | 3.93 |
| A3-highgpu-1g | 1 | 3.00 | |
| Azure | Standard-NC40ads-H100-v5 | 1 | 6.98 |
| Oracle | BM.GPU. H100.8 | 8 | 10.00 |
Appendix F Blocking Cost and One-time Offline Cost for CaRL-EM
We report the blocking cost per dataset, as shown in Table 9
| Dataset | Cost ($) |
|---|---|
| AG | |
| DA | |
| DS | |
| IM | |
| IT | |
| TT | |
| WA |
We report the token and GPU usage during training on AB, as shown in Table 10.
| Item | Time/h | Tokens | Cost ($) |
|---|---|---|---|
| Match | - | 335,592 | 0.04 |
| Compare | - | 1,694,532 | 0.18 |
| Select | - | 2,122,750 | 0.32 |
| PPO training (H100) | 2.75 | - | 16.45 |
We evaluate the overall end-to-end cost by breaking it down into three distinct parts: the initial blocking cost to build candidate sets, the model training expense, and the inference cost during deployment. Unlike supervised baselines like Ditto, which require separate training for every target dataset, CaRL-EM only needs to be trained once on the source dataset (AB). The learned controller then transfers zero-shot to new target datasets without any retraining. To provide a clear and transparent comparison, we report these three individual components alongside their total. We report the end-to-end cost, as shown in Table 7.
Training Scalability and Wall-Clock Analysis. To assess the practical scalability of PPO-based controller adaptation, we compare CaRL-EM and Ditto under different training set sizes on Walmart–Amazon (WA). We construct four subsets from WA′, containing 25%, 50%, 75%, and 100% of the training set, respectively. We fine-tune each method on these subsets, evaluate on the same WA test set used in the main paper, and record wall-clock time. We report fine-tuning rather than training from scratch because it is the more practical adaptation setting.
Figure 7 shows that CaRL-EM incurs higher wall-clock fine-tuning time, mainly because rollout collection repeatedly invokes backend operators. However, its downstream performance remains relatively stable across training sizes. By contrast, Ditto is much faster to fine-tune, but its performance degrades substantially in the low-data regime. In our setting, the controller itself is not the computational bottleneck: when the LLM operators are replaced with a simulator backend, tuning at the same scale (100k steps) takes only about 1.6 minutes, indicating that backend inference latency, rather than PPO optimization itself, dominates the overall wall-clock cost.