AI Contextual Measurement for Recovering Individual and Group-Level Effects:
Validation Against Survey Measures and an Occupational Application
Abstract
Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone.
We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets, not for replacing direct survey measurement or entire batteries of jointly missing items.
1 Introduction
Many social science questions involve variables that vary both between groups and within groups. Workers are nested within occupations, students within schools, employees within firms, patients within hospitals, and residents within neighborhoods. Researchers frequently want to know whether an association reflects differences among groups, differences among individuals within the same group, or both. This distinction is central to contextual analysis. If a characteristic is associated with an outcome, does the association arise because groups differ in their average characteristics, because individuals differ from others in the same group, or because both levels matter?
AI-based measurement offers new opportunities for studying constructs that are absent from conventional surveys. Large language models can score occupations, organizations, tasks, texts, or other entities on abstract dimensions such as autonomy, technological intensity, interpersonal orientation, or creativity. Many such measures are defined at the group or entity level. For example, an AI system may assign a single autonomy score to an occupation, firm, school, or neighborhood. Such scores can be useful for studying between-group differences, but they cannot determine whether individuals within the same group differ from one another or whether such within-group deviations matter for outcomes.
At the same time, recent work on AI-assisted survey prediction, LLM-based item imputation, and virtual survey respondents shows that language models can predict missing or unasked survey responses from respondent profiles, survey-question text, temporal information, or partial respondent attributes (Kim and Lee, 2023; Ji et al., 2024; Zhao et al., 2025). Our purpose is not to claim that respondent-level AI prediction is itself new. Instead, we ask whether AI-derived respondent-level measures can be used for a different inferential task: contextual measurement. The key question is whether an AI-derived variable can recover not only an observed survey measure, but also the individual-level and group-level associations that would be estimated if the survey measure were observed.
This paper proposes AICOME, AI COntextual MEasurement, as a framework for constructing, aggregating, and validating AI-derived measures for contextual analysis. Let denote the group to which individual belongs. A respondent-level AI measure can be decomposed into a group mean and an individual deviation from that group mean:
| (1) |
This decomposition turns AICOME into contextual measurement. It allows researchers to use both and , or equivalently and , in models that distinguish individual-level and group-level associations. Standard group-level AI scores can measure differences across groups, but they cannot identify within-group heterogeneity because their within-group deviation is mechanically zero.
The framework is easiest to describe using an underlying-concept notation. Let index an underlying concept of theoretical interest, such as computer use, technology use, management responsibility, or foreign-language use. Let be the true level of concept for individual . Let denote respondent ’s survey-based indicator of concept , where represents the survey-question wording, response scale, and measurement protocol. Let denote respondent ’s AI-derived indicator of the same concept, where is the respondent-specific feature vector supplied to the AI system and is the prompt protocol for concept . The outcome is denoted by , and regression controls are denoted by . This notation emphasizes that neither nor is treated as the true concept. They are alternative indicators of the same underlying concept, potentially shaped by different measurement protocols, response scales, and respondent-specific feature information. The AI measure is therefore not designed merely to reproduce the observed survey response. For example, a survey item may be binary while the AI prompt elicits a richer ordered score. AICOME asks whether the contextual structure obtained from supports substantive conclusions similar to those obtained from when validation data are available.
We illustrate the framework using job satisfaction and occupational groups in the 2022 China Family Panel Studies (CFPS). In the empirical application, is job satisfaction and is occupation. The setting is useful for two reasons. First, CFPS contains rich respondent and job information that can serve as features for AI measurement. Second, CFPS contains several observed work-related variables that allow validation against conventional survey measures . Occupations provide one illustration of the general grouping structure. The same framework is potentially applicable to other hierarchical settings, such as students within schools, employees within firms, patients within hospitals, and residents within neighborhoods.
The paper has three main contributions. First, we develop a validation framework for AI contextual measurement that goes beyond response-level correlation. Existing validation studies focus primarily on response recovery, item-level accuracy, or distributional similarity. We treat these as useful but insufficient evidence for contextual applications. Our primary validation target is recovery of contextual inference: whether AI-derived measures recover explanatory power, preserve qualitative coefficient directions, reproduce individual-level and group-level conclusions, and generate similar fitted values. Coefficient-vector correlations and RMSE are reported as descriptive diagnostics, not as strict requirements that the AI measure numerically mimic the survey measure.
Second, we show that respondent-level AI measures can recover contextual decompositions for several observed CFPS dimensions. This is the most direct evidence that the approach is useful for contextual analysis rather than merely item-level prediction. Weekly hours provides the clearest case: both AI prompting strategies reproduce the large negative between- and within-occupation associations observed in CFPS, while occupation-level measures miss the individual-level signal.
Third, we identify boundary conditions. Performance is strong when rich respondent information is available and a limited number of concepts are missing, but it deteriorates when information is sparse or when several related constructs are missing simultaneously. These findings clarify the intended use of the framework. AICOME is not a substitute for direct survey measurement and should not be interpreted as creating information absent from the observed data. It is most defensible when researchers possess rich existing respondent information and seek to recover a limited number of theoretically important but unmeasured concepts. The later application to autonomy, people-things, creative-routine, and technology illustrates this use case, but because those latent dimensions lack direct CFPS validation measures, those results are interpreted as an application of the validated framework rather than as direct proof of latent truth.
Figure 1 summarizes the central distinction between the present study and recent AI-assisted survey-prediction research. Existing work has primarily evaluated the ability of large language models to recover missing responses, respondent attributes, or population distributions. Our focus instead is whether respondent-level AI measures can support contextual analysis. Consequently, the principal validation target is not response recovery alone, but recovery of between- and within-occupation inferences.
2 Related Literature and Positioning
2.1 Contextual models and occupational measurement
A central concern in the contextual-model literature is distinguishing between-group and within-group associations (Blalock, 1984; Enders and Tofighi, 2007).Researchers studying neighborhoods, schools, firms, and occupations often seek to determine whether outcomes are driven by group environments or by differences among individuals within the same group. Estimating such models requires measures that vary at the individual level. Group-level measures can be useful for describing between-group differences, but by construction they cannot identify within-group deviations.
A large occupational-measurement literature has developed resources such as the Dictionary of Occupational Titles and O*NET to characterize jobs and occupations (Peterson et al., 2001; Autor et al., 2003; Acemoglu and Autor, 2011; Deming, 2017). These resources have been widely used to measure task content, skill requirements, technological exposure, social skills, and other features of work. Their central strength is that they provide systematic and comparable measures across occupations. Their central limitation for the present purpose is that they are usually occupation-level measures. If all workers in the same occupation receive the same score, the measure can identify only between-occupation variation.
Recent advances in large language models have created a related AI-measurement literature. Researchers have used language models to score occupations, estimate exposure to artificial intelligence, classify work tasks, generate synthetic respondents, and construct new social-science measures (Felten et al., 2021; Eloundou et al., 2024). These approaches overlap with the present paper because they use AI systems to generate quantities that were not directly measured in conventional datasets. However, most existing AI occupational measures remain group-level: occupations, occupational tasks, industries, firms, or texts receive scores, and individuals inherit the score attached to their group or label. Our contribution is to generate respondent-level measures that can be decomposed into occupation means and within-occupation deviations, making contextual analysis possible.
2.2 AI survey prediction, imputation, and synthetic respondents
Individual AI measurement is related to several emerging uses of large language models in simulating human behavior and in survey research. Horton, Filippas, and Manning (2023) (Horton et al., 2023) use LLMs as simulated economic agents to reproduce and explore behavioral responses in hypothetical experiments. Argyle et al. (2023) (Argyle et al., 2023) use LLMs to generate synthetic respondents (”silicon samples”) and evaluate whether the resulting responses reproduce patterns observed in human survey data. Another line of work uses LLMs to augment surveys, address Item Non-Response or as Virtual Survey Respondents (Kim and Lee, 2023, (Kim and Lee, 2023), Ji et al., 2024, (Ji et al., 2024), Zhao et al., 2025, (Zhao et al., 2025)).
These studies overlap with the present paper because all use AI systems to infer quantities not directly observed in a survey. The distinction is in the estimand and validation target. Prior work primarily asks whether AI systems can reproduce responses, opinions, attributes, or distributions. We ask whether respondent-level AI-derived measures can support contextual analysis. Specifically, we evaluate whether AI-generated variables recover explanatory power, preserve coefficient directions, and reproduce substantive between-within conclusions obtained from observed CFPS measures. The validation target is therefore not item-level prediction alone, but recovery of substantive contextual inferences.
This distinction is important because high response-level similarity need not imply valid downstream inference. Synthetic survey data may match marginal averages while producing different variation or regression coefficients, and results may vary with prompt wording or model changes (Bisbee et al., 2024). Motivated by this concern, our validation strategy compares not only correlations between AI-derived and observed measures, but also incremental explanatory power, coefficient directions, contextual conclusions, coefficient-vector similarity, and fitted-value similarity.
2.3 Construct validation, generated regressors, and measurement error
The validation problem in the present study is related to but not identical to either a classical generated-regressor problem or a simple measurement-error problem. In generated-regressor settings, a first-stage estimate is often assumed to converge to a well-defined target regressor, and the central issue is how first-stage estimation uncertainty affects second-stage standard errors (Pagan, 1984; Murphy and Topel, 1985) (Pagan, 1984; Murphy and Topel, 1985). See also Escanciano and P´erez-Izquierdo 2023 Escanciano and Pérez-Izquierdo (2023) for an automatic locally robust GMM framework for inference with machine-learning-generated regressors. Recently, Ludwig et al. (2026) Ludwig et al. (2026) develop an econometric framework for using LLM outputs in prediction and estimation, emphasizing that valid downstream inference for LLM-measured concepts requires validation data to debias and account for LLM errors. In the present setting, the AI-derived measure is not assumed to converge to the observed survey response as sample size grows. Nor do we assume that the observed survey response is a perfect measure of the underlying job concept that affects satisfaction. Both the observed survey variable and the AI-derived variable may be imperfect indicators of an underlying concept.
For this reason, the central validation question is whether the AI-derived measure yields substantively similar inferences to the observed survey measure when no perfect validation data are available. This framing is closer to construct validation than to a narrow two-step inference correction. The construct-validity tradition emphasizes that validity is not established by a single statistic, but by evidence that the operational measure supports the intended interpretation and use (Cronbach and Meehl, 1955; Messick, 1989; Adcock and Collier, 2001) (Cronbach and Meehl, 1955; Messick, 1989; Adcock and Collier, 2001). Here the intended use is contextual analysis. We therefore evaluate whether AI-derived measures reproduce the contextual inferences obtained from conventional survey measures.
A useful way to interpret these imperfect validation metrics is through an underlying-concept proxy framework. Let denote the unobserved concept described by the survey item or AI prompt, let denote any candidate indicator of that concept, such as a survey measure , a rich-prompt AI measure , or a survey-prompt AI measure , and let denote observed controls. The concept is unobserved in the dataset, but it is not an arbitrary abstraction: it is described in words by the measurement task itself, such as computer use, management responsibility, weekly hours, or foreign-language use.
Let denote the population linear prediction of from variables . We use the following linear proxy-validity assumption: once the outcome-relevant component associated with the underlying concept and controls has been extracted, the remaining outcome residual is uncorrelated with the candidate indicator,
| (2) |
Under this assumption, the prediction based on is the population linear projection of the oracle prediction based on onto the information available in . Consequently,
| (3) |
Thus, among competing indicators satisfying the same proxy-validity condition, the one with larger produces fitted values more highly correlated with the fitted values that would be obtained using the underlying concept itself.
The corresponding incremental version removes the part of the prediction already explained by the controls. Define
| (4) |
and let . Then
| (5) |
where the partial correlation is the ordinary correlation between and . This formulation is useful because it relates incremental explanatory power to the similarity between the incremental prediction signal from the candidate indicator and the incremental prediction signal from the underlying concept. In the scalar linear case, this reduces to the familiar attenuation-style identity
| (6) |
These identities do not require the survey variable to be the true concept, nor do they require AI and survey coefficients to be numerically equal. They also allow an AI-derived indicator to have greater explanatory power than a coarse survey item if it captures more outcome-relevant signal about the same concept. Individual coefficient signs, coefficient-vector correlations, and RMSE are useful diagnostics, but they are not guaranteed by the framework and are not the primary estimands. Favorable agreement in those quantities may suggest that the AI and survey indicators capture substantial common signal from the same underlying concept.
3 AICOME Framework
3.1 Group-level AI scores
A common measurement strategy is to assign a score directly to a group or entity. In an occupational application, this means assigning each occupation a score for a dimension such as technology or autonomy. In other settings, the group might be a firm, school, hospital, neighborhood, or region. If individual belongs to group and the AI score is constructed directly at the group level, the measure takes the form
| (7) |
This measure can be used to estimate how outcomes differ across groups. But because every member of the same group has the same value, the within-group deviation is zero:
| (8) |
Thus, group-level AI scores cannot test whether within-group heterogeneity exists. AICOME differs from direct group-level AI scoring by first constructing respondent-level measures and then deriving group-level aggregates from those respondent-level measurements.
3.2 Respondent-level AI measures
AICOME begins by constructing respondent-level indicators of an underlying concept. For , the level of concept on individual , the AI-derived measure is written as , where denotes respondent ’s feature vector for concept and denotes the prompting protocol. In the CFPS application, the feature set always includes group membership (if available), represented by occupation, but the remaining features may vary across target concepts. The rich-prompt strategy uses a richer version of questionnaire to describe the concept, while the survey-prompt strategy uses a more concise prompt structured around the questionnaire used in the actual CFPS survey. Both strategies yield respondent-level predicted scores for each dimension.
Because these scores vary within groups, they can be aggregated to group means and used in contextual models, paralleling standard centering logic in multilevel and contextual models (Enders and Tofighi, 2007). The group mean is
| (9) |
The key contextual model can be written either in the raw contextual parameterization as
| (10) |
or in the deviation contextual parametrization as
| (11) |
where is the outcome, is the respondent-level AI measure for dimension , is the corresponding group mean, and is a vector of controls. In the CFPS application, is job satisfaction and is occupation. We refer to as the individual or within-group association, as the contextual contrast, and as the group effect or between-group effect. The three regression coefficients satisfy the relation
| (12) |
When the ’s have the same sign, is smaller in magnitude and harder to detect than . This can also be aggravated by variance inflation: the standard error of is often larger than that of , since the group mean ’s are likely to be more associated linearly to the other right-hand side variables ’ in the raw parametrization than to the ’s in the deviation parametrization.
All reported regression coefficients are standardized: and components are standardized by their respective standard deviations, and the raw , its group mean and deviation from group mean are all standardized by the standard deviation of the raw ’s so that holds. Standard errors are cluster-robust at the occupation level.
The controls in the CFPS application include age, gender, education and marital status from the last interview, urban residence, agricultural hukou, public or state-owned employment, whether the job is outdoors, health, and log income. Weekly hours is excluded from the control set when weekly hours is a focal validated dimension, and included otherwise.
3.3 Validation framework
The empirical validation proceeds at four levels. First, response-level validation examines correlations between AI-derived measures and observed survey measures . These correlations may suggest shared concept-related signal, but are not the final validation target. Second, model-level validation asks whether AI-derived measures recover incremental explanatory power and coefficient directions in non-contextual regressions. Third, contextual validation asks whether AI-derived measures recover the individual-level and group-level conclusions obtained from survey measures. Fourth, boundary validation examines when the approach deteriorates, using reduced-information and simultaneous-missingness designs.
4 Data and Measures
4.1 Data
The empirical illustration uses the 2022 wave of the China Family Panel Studies (CFPS). The data are from China Family Panel Studies (CFPS), funded by 985 Program of Peking University and carried out by the Institute of Social Science Survey of Peking University. CFPS is useful for this validation exercise because it contains rich respondent information, occupational membership, job-related characteristics, and subjective well-being measures. These features make it possible to construct respondent-level AI measures, aggregate them within occupational groups, and validate the resulting contextual structure against directly observed survey measures.
The outcome variable is job satisfaction. In the present application, the grouping variable is occupation, identified using cleaned occupational codes. Occupations are used as the empirical contextual unit for constructing group-level aggregates and estimating contextual models. The methodological framework, however, is not occupation-specific. The same logic applies to other settings in which individuals are nested within larger social units.
The validation analysis uses four observed work-related CFPS dimensions: computer use, foreign-language use, weekly hours, and management responsibilities. In the notation of the framework, these observed survey variables are denoted by . They provide benchmarks for evaluating AI-derived measures . The later latent application examines autonomy, people-things, creative-routine, and technology, which are not directly observed in CFPS and therefore cannot be validated in the same way.
Detailed variable definitions, cleaning rules, AI feature sets , and prompt protocols will be documented in the appendices.
4.2 AI-derived measures
For each observed CFPS dimension, we compare the survey measure , the rich-prompt AI measure , and the survey-prompt AI measure . We also compare individual AI measures with occupation-mean versions of those measures, occupation-level AI scores, and embedding-based occupation scores in non-contextual validation models. For contextual validation, occupation-level scores cannot identify within-occupation deviations, so the contextual comparisons focus on survey respondent-level measures, rich-prompt individual AI measures, and survey-prompt individual AI measures.
For the latent application, we examine autonomy, people-things, creative-routine, and technology. These dimensions are not directly observed in CFPS. Therefore, the latent analysis evaluates robustness across rich-prompt and survey-prompt AI measures rather than direct truth recovery. It also serves an in illustration how AICOME can be applied to do contextual analysis for a concept that is not included in an actual survey.
5 Validation Using CFPS-Measured Dimensions
5.1 Correlation validation of AI-derived measures
Before examining contextual models, we first assess the correspondence between the AI-derived measures and the observed CFPS variables used for validation. Table 1 reports the correlations between the true survey measures and the corresponding rich-prompt and survey-prompt AI measures. Since later in the raw or deviation parametrization of the contextual models, each variable may appear on the right-hand side in three different forms, the raw , its group mean and deviation from group mean , we include the correlations to the corresponding CFPS variable also in three forms.
| Rich AI vs. CFPS | Survey AI vs. CFPS | |||||
|---|---|---|---|---|---|---|
| Construct | vs. | vs. | vs. | vs. | vs. | vs. |
| Computer Use | 0.881 | 0.944 | 0.844 | 0.935 | 0.930 | 0.896 |
| Foreign Language | 0.868 | 0.887 | 0.869 | 0.696 | 0.814 | 0.684 |
| Weekly Hours | 0.879 | 0.976 | 0.853 | 0.915 | 0.986 | 0.896 |
| Management | 0.981 | 0.912 | 0.971 | 0.971 | 0.959 | 0.965 |
The correlations are generally high, ranging from 0.684 to 0.986. Management exhibits particularly strong agreement, with correlations of 0.981 for the rich-prompt measure and 0.971 for the survey-prompt measure in the raw form. Computer use and weekly hours also show strong correspondence under both prompting strategies. Foreign language is the most challenging construct, particularly for the survey-prompt measure, which attains a correlation of 0.684 with the observed CFPS variable in the deviation form.
These results provide a useful preliminary validation benchmark. In the common-concept interpretation, high correlations may indicate that the AI and survey indicators share substantial concept-related signal.
5.2 Sensitivity to available respondent information
The high validation correlations reported in Table 1 naturally raise a question about the information available to the AI model. AICOME does not generate information from nothing. Rather, it constructs indicators of an underlying concept from respondent characteristics supplied in the prompt. Consequently, the usefulness of AICOME depends on the amount of concept-relevant information contained in the available covariates.
To evaluate this dependence, we compare two information sets. The first is the full-information specification used throughout the main analysis. The second is a reduced-information specification intended to mimic a shorter survey containing only occupation and basic demographic characteristics. The reduced information set is
This specification excludes all job-specific information.
The full information set varies by target dimension and is intentionally designed to use information that would typically be available in a labor-force or household survey. The common variables used for computer use, foreign-language use, weekly hours, and management responsibilities include occupation, age, employer type, education in the last interview, organization size, and urban residence. For computer use, the AI additionally receives work location, employer type, management responsibilities, direct reports, and number of subordinates. For foreign-language use, the AI additionally receives province, management responsibilities, number of subordinates, gender, and marital status from the last interview. For weekly hours, the AI additionally receives industry, income, management responsibilities, promotion history, promotion expectations, tenure, contract status, night-shift frequency, weekend work, and on-call duties. For management responsibilities, the AI additionally receives industry, income, promotion history, promotion expectations, party membership, tenure, contract status, and marital status from the last interview.
Importantly, these full-information specifications are not constructed from variables that are inherently unavailable in survey settings. Most of the variables—such as occupation, industry, education, age, gender, income, job tenure, employer type, supervisory responsibilities, organizational size, and working hours—are routinely collected in labor-force and household surveys, while more detailed employment characteristics, such as work schedules and promotion history, are available in more detailed surveys. The full-information specification therefore represents a rich but empirically realistic survey environment in which one survey item is missing and must be inferred from other observed respondent characteristics.
| Rich Prompt | Survey Prompt | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| vs. | vs. | vs. | vs. | vs. | vs. | |||||||
| Dimension | Full | Reduced | Full | Reduced | Full | Reduced | Full | Reduced | Full | Reduced | Full | Reduced |
| Computer Use | 0.881 | 0.555 | 0.944 | 0.913 | 0.844 | 0.170 | 0.935 | 0.385 | 0.930 | 0.832 | 0.896 | 0.103 |
| Foreign Language | 0.868 | 0.199 | 0.887 | 0.571 | 0.869 | 0.075 | 0.696 | 0.221 | 0.814 | 0.651 | 0.684 | 0.091 |
| Weekly Hours | 0.879 | -0.019 | 0.976 | -0.216 | 0.853 | 0.065 | 0.915 | 0.028 | 0.986 | -0.077 | 0.896 | 0.065 |
| Management | 0.981 | 0.372 | 0.912 | 0.857 | 0.971 | 0.046 | 0.971 | 0.369 | 0.959 | 0.830 | 0.965 | 0.070 |
| Average | 0.902 | 0.277 | 0.930 | 0.531 | 0.884 | 0.089 | 0.879 | 0.251 | 0.922 | 0.559 | 0.860 | 0.082 |
The differences are substantial. In the raw form, under the rich-prompt strategy, the average correlation across the four validation dimensions declines from 0.902 to 0.277. Under the survey-prompt strategy, the average correlation declines from 0.879 to 0.251. The deterioration is particularly pronounced for weekly hours and foreign-language use. For weekly hours, the correlation falls from approximately 0.9 to essentially zero. Similar collapses occur for foreign-language use. Computer use and management remain positively correlated with the true survey measures but exhibit large declines relative to the full-information specifications.
These results indicate that AICOME depends critically on the availability of informative respondent characteristics. When only occupation and basic demographic information are available, validation performance deteriorates sharply. Conversely, when richer job-related information is available, AI-derived measures closely track the corresponding observed survey responses. This result distinguishes the present framework from stronger claims about synthetic replacement of survey data. The AI model performs well only when informative respondent and job characteristics are available; when the available information is sparse, performance deteriorates sharply. The method therefore should not be interpreted as creating respondent-level information absent from the observed data.
This boundary condition is consistent with recent work on augmenting survey data with generative AI. Brynjolfsson et al. (2026) find that supplying LLMs with rich contextual information beyond demographics substantially improves predictive accuracy in economic survey applications, whereas changes to prompting strategy alone yield smaller gains. Our results extend this pattern to a contextual-analysis setting: richer respondent information improves not only response-level correspondence, but also the ability to recover between-occupation and within-occupation structure.
5.3 Single-dimension non-contextual validation
Correlation alone is not the primary criterion for evaluating the usefulness of the AI-derived measures. The remainder of the validation section focuses on a more demanding test: whether AI-derived measures reproduce explanatory power, contextual conclusions, and multivariate regression results obtained from the observed CFPS variables.
We begin with non-contextual models that add one measured job dimension at a time. This first validation exercise compares all available measurement strategies, including individual AI measures, occupation means, occupation-level AI scores, and embedding-based scores. Table 3 reports values. We emphasize the gain in relative to the controls-only model because, under the underlying-concept interpretation, this gain measures the amount of outcome-relevant concept signal retained by the indicator.
| Measure | Computer | Foreign language | Weekly hours | Management |
|---|---|---|---|---|
| Controls only | 0.0759 | 0.0759 | 0.0541 | 0.0756 |
| CFPS individual measure | 0.0787 | 0.0763 | 0.0734 | 0.0798 |
| CFPS occupation mean | 0.0788 | 0.0767 | 0.0600 | 0.0782 |
| Rich individual AI | 0.0780 | 0.0761 | 0.0735 | 0.0805 |
| Rich AI occupation mean | 0.0774 | 0.0765 | 0.0608 | 0.0782 |
| Survey-prompt individual AI | 0.0795 | 0.0765 | 0.0719 | 0.0806 |
| Survey-prompt AI occupation mean | 0.0787 | 0.0778 | 0.0606 | 0.0782 |
| Occupation-level AI score | 0.0769 | 0.0801 | 0.0542 | 0.0783 |
| Embedding occupation score | 0.0770 | 0.0762 | 0.0559 | 0.0784 |
The clearest case is weekly hours. The CFPS individual weekly-hours measure raises from 0.0541 to 0.0734. The rich and survey individual AI measures produce very similar values, 0.0735 and 0.0719. By contrast, the CFPS occupation mean produces a smaller of 0.0600, and the occupation-level AI score contributes essentially no explanatory power beyond controls. This pattern illustrates why individual-level measurement can matter: when a construct has substantial within-occupation signal, occupation-level measures cannot recover the full association.
In addition, the table illustrates why the CFPS survey measure should not be treated automatically as the highest-performing indicator of the underlying concept. Because survey items may be coarse and AI measures may use richer feature information or response scales, an AI-derived indicator can in some cases have comparable or greater explanatory power than the survey indicator while still measuring the same concept.
5.4 Single-dimension contextual validation
The next validation step asks whether AI-derived individual measures recover the contextual decomposition obtained from observed CFPS measures. For each construct, we estimate the equivalent raw contextual parameterization:
| (13) |
or the deviation contextual parametrization:
| (14) |
where is the individual measure and is the corresponding occupation mean. In these parameterizations, is the individual or within-occupation coefficient and is the contextual contrast associated with the occupation mean; the group effect or between-occupation coefficient is . Table 4 reports the results.
| Construct | Method | Individual (SE) | Contextual (SE) | Group (SE) | |
|---|---|---|---|---|---|
| Computer use | CFPS | 0.0799 | 0.046 (0.020) | 0.084 (0.031) | 0.130 (0.031) |
| Computer use | Rich AI | 0.0782 | 0.052 (0.027) | 0.031 (0.030) | 0.084 (0.025) |
| Computer use | Survey AI | 0.0803 | 0.061 (0.021) | 0.064 (0.031) | 0.125 (0.030) |
| Foreign language | CFPS | 0.0769 | 0.012 (0.013) | 0.070 (0.038) | 0.081 (0.031) |
| Foreign language | Rich AI | 0.0765 | 0.006 (0.014) | 0.048 (0.028) | 0.053 (0.021) |
| Foreign language | Survey AI | 0.0779 | 0.011 (0.017) | 0.088 (0.031) | 0.099 (0.027) |
| Weekly hours | CFPS | 0.0749 | -0.135 (0.016) | -0.112 (0.037) | -0.247 (0.039) |
| Weekly hours | Rich AI | 0.0752 | -0.136 (0.017) | -0.120 (0.037) | -0.256 (0.037) |
| Weekly hours | Survey AI | 0.0738 | -0.129 (0.016) | -0.128 (0.038) | -0.257 (0.039) |
| Management | CFPS | 0.0806 | 0.055 (0.012) | 0.072 (0.028) | 0.126 (0.029) |
| Management | Rich AI | 0.081 | 0.061 (0.012) | 0.052 (0.025) | 0.113 (0.027) |
| Management | Survey AI | 0.081 | 0.063 (0.011) | 0.047 (0.025) | 0.110 (0.027) |
The contextual validation results show that AI-derived measures largely reproduce the observed-variable decomposition. Weekly hours is the strongest case. The CFPS individual effect is , and the rich and survey AI versions estimate and . The contextual contrasts are also significantly negative in all three specifications. These demonstrate that the method can detect a large individual effect and a large contextual effect when such signals exist.
Management is another strong validation case. All three methods estimate positive individual effects and positive occupation group effects. However, the contextual effects are weaker and not universally significant. Computer use shows broadly consistent positive effects, though the rich-prompt method attenuates the contextual effect. Foreign language shows a weak individual effect and a stronger occupation-level component across all three approaches, though the survey-prompt occupation-mean contrast is larger than the observed or rich-prompt estimate. Overall, the AI measures recover the qualitative contextual story for all four observed dimensions. This agreement is not mechanically guaranteed by the framework. Rather, it is favorable empirical evidence that the AI and survey indicators may capture substantial common signal from the same underlying concepts. Because individual and occupation-level components may be attenuated differently across proxies, we interpret sign concordance, , fitted values, and substantive conclusions as the primary evidence rather than requiring exact coefficient equality.
5.5 Joint contextual validation under simultaneous missingness
The preceding validation exercises consider one survey dimension at a time. In those settings, a missing construct can be inferred using other observed respondent and job-characteristic information. Many practical applications, however, involve multiple missing dimensions simultaneously. In such settings, computer use cannot be used to predict management responsibilities, management responsibilities cannot be used to predict weekly hours, and so forth, because all of these dimensions are themselves unavailable.
To evaluate this more demanding scenario, we estimate a joint contextual model in which all four validated dimensions—computer use, foreign-language use, weekly hours, and management responsibilities—are treated as simultaneously unobserved. We therefore impose a joint exclusion rule. Any variable excluded when validating one dimension is excluded from the prompts used to generate all four AI measures. In the present application, this removes direct or closely related measures of computer use, foreign-language use, weekly hours, management responsibilities, supervisory status, and number of subordinates from all four AI imputations.
This specification is substantially more challenging than the single-dimension validation exercises. After applying the joint exclusion rule, the information common to all four AI prompts consists primarily of occupation, education degree, employer type, age, organizational size, and urban residence. Some additional dimension-specific information remains available when it is not part of the union of exclusions. This design more closely resembles a realistic survey setting in which several job-characteristic measures were never collected.
This exercise is conceptually related to recent work evaluating LLM-based survey imputation under different missingness mechanisms. Holtdirk et al. (2026) study in-context learning for imputing public-opinion survey data across MCAR, MAR, and MNAR designs. Our design addresses a different validation target: several substantively connected job-characteristic variables are jointly unavailable for every respondent, and the question is whether the resulting AI-derived measures recover contextual coefficient structure rather than individual item responses alone.
| Effect | CFPS | Rich AI4 | Survey AI4 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| (SE) | (SE) | (SE) | (SE) | (SE) | (SE) | (SE) | (SE) | (SE) | |
| Computer | 0.094 (0.037) | 0.053 (0.042) | 0.041 (0.023) | 0.053 (0.043) | 0.043 (0.050) | 0.010 (0.040) | 0.026 (0.041) | 0.037 (0.047) | -0.011 (0.039) |
| Foreign language | 0.043 (0.029) | 0.026 (0.031) | 0.017 (0.013) | -0.020 (0.035) | -0.005 (0.041) | -0.015 (0.018) | 0.025 (0.034) | -0.025 (0.035) | 0.050 (0.020) |
| Hours | -0.137 (0.048) | 0.001 (0.049) | -0.138 (0.016) | -0.153 (0.047) | -0.059 (0.050) | -0.094 (0.017) | -0.192 (0.041) | -0.103 (0.046) | -0.089 (0.018) |
| Management | 0.067 (0.025) | 0.011 (0.025) | 0.055 (0.011) | 0.065 (0.017) | 0.010 (0.029) | 0.054 (0.023) | 0.069 (0.016) | 0.063 (0.028) | 0.006 (0.026) |
| Controls-only | 0.0517 | 0.0517 | 0.0517 | ||||||
| Full | 0.0786 | 0.0659 | 0.0684 | ||||||
| 0.0269 | 0.0142 | 0.0167 | |||||||
| 5561 | 5561 | 5561 | |||||||
| Occupation clusters | 298 | 298 | 298 | ||||||
- •
Note: , , and denote the group, contextual, and individual effects, respectively.
Table 5 reports the results. Relative to the observed CFPS measures, the AI specifications exhibit lower incremental explanatory power and attenuation or instability in several smaller coefficients. The increase in from adding the eight contextual terms is 0.0269 using the observed CFPS measures, compared with 0.0142 for the rich-prompt AI measures and 0.0167 for the survey-prompt AI measures. In the vector interpretation, this reduction indicates that the AI4 measures retain less of the outcome-relevant concept signal than the observed survey indicators under simultaneous missingness.
No contextual effect is found to be significant in this joint model. Despite this attenuation, the overall qualitative pattern remains recognizable for the strongest signals. Weekly hours retains both negative occupation group effect and negative individual within-occupation associations. Management remains positively associated with job satisfaction. However, several smaller effects become attenuated, less precise, or less stable across prompting strategies. Thus, the AI measures recover some broad features of the contextual structure, especially for the stronger hours and management signals, but the joint-missingness design also shows that AICOME becomes less reliable when several related concepts are simultaneously unavailable.
| Similarity measure | Rich AI4 vs CFPS | Survey AI4 vs CFPS | Rich AI4 vs Survey AI4 | |||
|---|---|---|---|---|---|---|
| Correlation of eight focal coefficients | 0.939 | 0.869 | 0.878 | 0.579 | 0.884 | 0.740 |
| RMSE of eight focal coefficients | 0.035 | 0.033 | 0.045 | 0.056 | 0.038 | 0.039 |
| Mean absolute difference of eight focal coefficients | 0.029 | 0.026 | 0.041 | 0.051 | 0.032 | 0.033 |
| Correlation of z-component predicted satisfaction | 0.646 | 0.645 | 0.907 | |||
| RMSE of z-component predicted satisfaction | 0.151 | 0.156 | 0.072 | |||
| Correlation of full fitted satisfaction | 0.850 | 0.842 | 0.966 | |||
| RMSE of full fitted satisfaction | 0.149 | 0.153 | 0.068 | |||
- •
Note: Within each method comparison, the column summarizes agreement in the between-group and individual effects, whereas the column summarizes agreement in the contextual and individual effects. Prediction statistics are identical across the two parameterizations.
Table 6 summarizes the similarity of the joint contextual results. In the deviation parametrization, the correlations of the eight focal coefficient vectors are 0.939 for rich-prompt AI versus observed CFPS and 0.878 for survey-prompt AI versus observed CFPS. In the raw parametrization, the correlations of the eight focal coefficient vectors are 0.869 for rich-prompt AI versus observed CFPS and 0.579 for survey-prompt AI versus observed CFPS. These correlations indicate that the broad coefficient structure is still partially recovered, especially in the deviation parametrization . The RMSE and mean absolute difference statistics are useful descriptive diagnostics, but they are not strict validation targets because different indicators of the same concept may be attenuated by different amounts. The fitted-value correlations provide a complementary summary of whether the AI measures recover the outcome-relevant component of the contextual model.
The results suggest an important practical distinction. AICOME performs best when a missing concept can be inferred from other observed respondent characteristics. Recovering several related dimensions simultaneously is substantially more difficult because the dimensions can no longer be used to predict one another. Consequently, AICOME should be viewed primarily as a tool for recovering limited numbers of missing constructs rather than an exact replacement for jointly missing batteries of related survey questions.
6 Application to Latent Dimensions
The latent-dimension application should be interpreted differently from the CFPS validation analysis. For autonomy, people-things, creative-routine, and technology, CFPS does not contain direct validation measures. The analysis therefore cannot establish that the AI-derived variables recover the true latent constructs. Instead, it asks whether two independently designed prompting strategies yield similar contextual conclusions and whether the resulting associations are substantively plausible. The validated CFPS dimensions provide the methodological proof-of-concept; the latent dimensions illustrate how the framework can be applied when direct validation is unavailable.
6.1 Single-Dimension Latent Analyses: Diagnosis and Repair
Because the four occupational dimensions considered here are latent constructs without directly corresponding CFPS survey measures, their evaluation requires combining outcome-based evidence with substantive validation. We therefore examined each AI-generated dimension separately in a contextual model of job satisfaction and inspected the occupations receiving high and low average scores. This analysis revealed unexpected results for Autonomy. We first describe the diagnosis and revision of this measure and then summarize the final single-dimension findings for all four dimensions.
6.1.1 Autonomy: Initial evidence of a measurement problem
For each latent dimension , we estimated two algebraically equivalent parameterizations: the deviation-contextual model parameterized by and the raw-contextual model parameterized by , where , is the group (or the between group) effect, is the individual (or within group) effect, and is the contextual effect. The coefficient of in the deviation-contextual model supplies , the coefficient of in the raw-contextual model supplies , and the coefficient of , or equivalently , supplies .
Table 7A reports the initial results for Autonomy using the full feature set , which included the time-flexibility variable qg604.
| Dimension | Prompt | (SE) | (SE) | (SE) | Baseline | Full |
|---|---|---|---|---|---|---|
| Autonomy | Rich | (0.046) | (0.046) | 0.003 (0.015) | 0.0740 | 0.0740 |
| Autonomy | Survey | 0.005 (0.046) | (0.047) | 0.007 (0.015) | 0.0740 | 0.0740 |
- •
Note: is the coefficient of in the deviation-contextual model; is the coefficient of in the raw-contextual model; and is the coefficient of or , respectively. Standard errors are clustered by occupation. AI scores were constructed from the full feature set , including time flexibility qg604.
The initial Autonomy results were particularly anomalous. The group, contextual, and individual estimates were all close to zero under both prompt designs, and adding the AI measure produced no detectable improvement in . These results did not by themselves prove that the measures were invalid, but they were sufficiently weak to motivate further examination of their construction.
6.1.2 Diagnosing the role of time flexibility
Inspection of the feature subset used to construct the Autonomy measure led us to detect unusually strong dependence on time flexibility. With the original full feature set, the correlation between qg604 and AI-derived Autonomy was 0.955 under the rich prompt and 0.898 under the survey prompt. This time flexibility variable was also used in constructing the Creative–Routine measure. The corresponding correlations for Creative–Routine were 0.491 and 0.566. Thus, Autonomy in particular was behaving almost as a transformation of the time-flexibility response rather than as a measure drawing more evenly on the available respondent and occupational information.
This concentration would not necessarily be problematic if qg604 unambiguously measured desirable scheduling autonomy. The CFPS data, however, suggested a more complicated interpretation. Mean job satisfaction was 3.714 for , 3.807 for , and 3.681 for . The original qg604 code decreases with time flexibility. Thus, the category coded as having the greatest flexibility did not have the greatest mean job satisfaction, and the overall correlation between time flexibility and satisfaction was approximately zero.
Further descriptive investigation showed that 42.7% of respondents in the most flexible category were self-employed or worked in a family business, compared with 6.4% among respondents in the other two categories. The most flexible group was also less likely to have an employment contract, cash or material benefits, or work-related insurance; more likely to work every weekend or be on call; and had lower mean annual work income. These comparisons indicate that, in the CFPS setting, high time flexibility can accompany self-employment, irregular scheduling, and limited employment protection rather than representing conventional employee autonomy alone.
Occupation-level rankings were consistent with this concern. Under the original specification, the upper end of the Autonomy ranking (based on occupation mean) included pump operators and metal-craft production workers alongside organizational leaders, musicians, and numerous self-employed occupations. The rich-prompt ranking was likewise headed by musicians and several types of self-employed proprietors but also included funeral workers, pump operators, and quartz-glass production workers. The rankings were therefore not uniformly implausible, but they do not appear to be consistently capturing the occupational level signals on Autonomy in the usual sense.
Several alternative specifications were considered. The final revision removed qg604 from the feature subsets used to generate Autonomy and Creative–Routine. We denote this revised feature set by . This modification sharply reduced the residual association with time flexibility: the correlations fell to 0.250 and 0.213 for rich- and survey-prompt Autonomy, and to 0.116 and 0.098 for rich- and survey-prompt Creative–Routine.
6.1.3 Autonomy: Results after measurement repair
Table 7B presents the single-dimension models after reconstructing Autonomy using .
| Dimension | Prompt | (SE) | (SE) | (SE) | Baseline | Full |
|---|---|---|---|---|---|---|
| Autonomy | Rich | 0.062 (0.028) | 0.016 (0.028) | 0.047 (0.015) | 0.0740 | 0.0765 |
| Autonomy | Survey | 0.054 (0.026) | 0.000 (0.029) | 0.054 (0.015) | 0.0740 | 0.0763 |
- •
Note: Definitions follow Table 7A. Autonomy scores were reconstructed using , which excludes qg604. Minor discrepancies in reflect rounding.
The improvement for Autonomy was substantial and internally coherent. Under both prompt designs, positive group effects were detected, with estimates of 0.062 and 0.054. Positive individual effects were also detected, with estimates of 0.047 and 0.054. In contrast, the contextual estimates were 0.016 and 0.000 and were not distinguishable from zero. The revised results therefore indicate that occupations with higher mean Autonomy tend to have higher mean satisfaction and that respondents assigned higher Autonomy scores within occupations tend to report greater satisfaction. They do not provide evidence that occupational mean Autonomy has an additional contextual association after controlling for a respondent’s own score. Incremental increased from zero to approximately 0.0024 to 0.0025.
The diagnostic exercise illustrates a useful feature of the AICOME workflow. The initial outcome analysis did not simply provide an unfavorable substantive result. It revealed a measurement anomaly. Examining feature dependence and occupational rankings showed that the AI procedure had placed disproportionate weight on a variable whose empirical meaning differed from the intended construct. Removing that variable recovered coherent group and individual Autonomy effects without creating a contextual effect. All subsequent analyses therefore use , which omits the time flexibility variable previously used to construct AI measures for Autonomy and Creative–Routine.
6.1.4 Face validity for AI measures
We now summarize the AI measures constructed from regarding their face validity of the ranked occupation means for all 4 latent directions. This is important to know before further analyses, since there is no CFPS survey on these latent directions that can be used to check on their correlations.
We find that the face validity is generally reasonable even though not perfect. Occasional anomalies exist. For example, for survey prompt in the People–Things direction, one of the top-ranking occupations is gardening technicians. Upon further investigation, we find that this is an occupation with one individual, who has a management role with direct reports. The AI seems to have focused on these individual-level features to assign a high score on People–Things. Despite the occasional anomalies, the general pattern still clearly shows that AI is capable of producing a reasonable ranking on each direction, so that it is hard to confuse one direction with another just by looking at the rankings.
The revised Autonomy rankings showed greater substantive coherence compared to the initial ones. Self-employed business owners remained highly ranked, appropriately reflecting one form of work independence, but the upper portion of the distribution also included various kinds of artists, judges, researchers, professional specialists, organizational leaders, and religious professionals. The lower end included assembly-line workers, packers, simple physical laborers, sanitation workers, and workers in tightly structured manufacturing and production occupations. Although such rankings cannot establish criterion validity, the contrast is consistent with an interpretation based on occupational discretion rather than time flexibility alone as before.
For Creative–Routine, the AI measurement also produced clear occupational face validity. Actors, fine-art professionals, musicians, photographers, crafts and arts workers, journalists, editors, university teachers, researchers, and designers appeared near the creative end. Assembly-line workers, packers, simple physical laborers, postal workers, and routine production workers appeared near the routine end. This ordering was evident under both prompt designs.
For the occupational rankings on People–Things, teachers, nurses, physicians, childcare and domestic-service workers, sales and dining-service workers, translators, actors, and other communication- or service-intensive occupations tended to appear toward the people-oriented end. Machine and equipment operators, assemblers, miners, construction and production workers, and other occupations centered on physical objects or machinery tended to appear closer to the things-oriented end. The exact rank ordering varies between prompts, but the broader people-versus-things contrast is visible in both.
The Technology rankings also exhibited substantial face validity. Electronic, computer, electrical, aerospace, aviation, communication, and other engineering occupations, together with scientific and medical researchers, appeared near the high-technology end under the rich and survey prompts. Sanitation workers, simple physical laborers, packers, agricultural workers, street vendors, domestic-service workers, and several routine production occupations appeared near the low-technology end.
6.2 Final Single-Dimension AICOME Results
Tables 7C constitute the final preferred single- results for the four latent dimensions.
| Dimension | Prompt | (SE) | (SE) | (SE) | Baseline | Full |
|---|---|---|---|---|---|---|
| Autonomy | Rich | 0.062 (0.028) | 0.016 (0.028) | 0.047 (0.015) | 0.0740 | 0.0765 |
| Autonomy | Survey | 0.054 (0.026) | 0.000 (0.029) | 0.054 (0.015) | 0.0740 | 0.0763 |
| Creative–Routine | Rich | 0.043 (0.023) | 0.089 (0.032) | (0.026) | 0.0740 | 0.0757 |
| Creative–Routine | Survey | 0.043 (0.023) | 0.057 (0.028) | (0.021) | 0.0740 | 0.0753 |
| People–Things | Rich | 0.100 (0.019) | 0.015 (0.036) | 0.085 (0.026) | 0.0740 | 0.0819 |
| People–Things | Survey | 0.104 (0.019) | 0.054 (0.030) | 0.051 (0.022) | 0.0740 | 0.0808 |
| Technology | Rich | 0.045 (0.027) | 0.013 (0.029) | 0.032 (0.025) | 0.0740 | 0.0747 |
| Technology | Survey | 0.104 (0.030) | 0.057 (0.030) | 0.047 (0.019) | 0.0740 | 0.0770 |
6.2.1 Creative–Routine
The regression evidence for Creative–Routine was mixed. The group coefficient was positive but only modest relative to its standard error under both prompts. The rich-prompt contextual coefficient was positive, while the individual coefficient was negative. The survey-prompt version showed the same directional decomposition but weaker estimates. Because
| (15) |
a modest positive group coefficient combined with a negative individual coefficient mechanically produces a larger positive contextual estimate, which means that for two people with the same job creativity, the one with higher occupational creativity is more satisfied. However, this suppression-like pattern, together with the small , makes it difficult to interpret the Creative–Routine result as conclusive evidence of a contextual mechanism.
6.2.2 People–Things
People–Things and Technology are unchanged by the NoTimeFlex revision because their feature subsets did not include qg604.
People–Things produced the strongest and most consistent single-dimension result. Its group effect was 0.100 under the rich prompt and 0.104 under the survey prompt, with standard errors of 0.019 in both cases. Positive individual effects were also detected under both prompts. By contrast, the contextual estimates were substantially smaller and less precise, particularly for the rich measure. People–Things also produced the largest incremental , 0.0079 and 0.0069. The consistent interpretation is therefore a positive group association together with a positive individual association, but no conclusive evidence of an additional contextual effect.
6.2.3 Technology
The relationship of Technology with job satisfaction was less consistent across prompt designs. Under the survey prompt, the group coefficient was 0.104 (SE 0.030) and the individual coefficient was 0.047 (SE 0.019), while the contextual estimate was 0.057 (SE 0.030). Under the rich prompt, all three estimated effects were smaller relative to their standard errors, and was only 0.0008. The evidence therefore supports group and individual Technology effects under the survey-prompt specification, but these effects are not robustly reproduced by the rich-prompt measure. Neither prompt provides clear, consistent evidence of a distinct contextual effect.
6.2.4 Summary across the four single-dimension analyses
Several conclusions emerge from the preferred single- analyses.
First, group effects are the most consistently detected component. People–Things and revised Autonomy display positive group effects under both prompt designs. Technology displays a positive group effect under the survey prompt but not clearly under the rich prompt. Creative–Routine has a modest positive group estimate, but the evidence is weaker.
Second, individual effects are clearly detected for People–Things and revised Autonomy. In both cases, the group association is accompanied by a positive individual association, while the contextual component is small or imprecise. The most defensible interpretation is therefore that the observed group differences largely reflect corresponding individual-level associations rather than an additional contextual mechanism.
Third, no contextual effect is established conclusively across prompt designs. The strongest apparent contextual coefficients arise for Creative–Routine, but they coexist with negative individual estimates, weak group effects, and small gains in model fit. This internally conflicting decomposition prevents a confident contextual interpretation. The survey-prompt contextual estimates for People–Things and Technology are suggestive, but they are not consistently reproduced by the rich-prompt measures.
Finally, the dimensions differ in the strength and consistency of their outcome validity. People–Things provides the strongest evidence, with clear group and individual effects and the largest improvement in model fit. Revised Autonomy provides consistent evidence of group and individual effects after correcting the measurement problem. Technology has prompt-dependent regression evidence. Creative–Routine has its group, contextual, and individual estimates point in different directions and remain difficult to interpret.
Overall, the single-dimension analyses demonstrate both the promise and the diagnostic value of AICOME. They recover substantively recognizable occupational constructs and detect several group and individual associations with job satisfaction. At the same time, they show why gross group, contextual, and individual components must be distinguished carefully: a positive group association does not itself establish a contextual effect, and a large contextual coefficient can arise from opposing group and individual components. The next subsection evaluates whether these conclusions persist when the four latent dimensions are included jointly.
6.3 Joint Four-Dimension Latent Analysis
The single-dimension analyses evaluate each latent occupational characteristic separately. Because Autonomy, People–Things, Creative–Routine, and Technology describe related aspects of work, however, their single-dimension associations may reflect variation shared across dimensions. We therefore estimated joint models including all four latent dimensions simultaneously. The preferred measures exclude time flexibility from the feature sets previously used to construct Autonomy and Creative–Routine, as described in the preceding diagnostic analysis.
As before, we report three effects for each dimension. The group effect, , is the coefficient of the occupational mean in the deviation-contextual parameterization. The contextual effect, , is the coefficient of the occupational mean in the raw-contextual parameterization. The individual effect, , is the coefficient of the respondent-level deviation or, equivalently, the respondent-level raw score. These coefficients satisfy
| (16) |
Table 8 presents the resulting joint decomposition.
| Dimension | Prompt | (SE) | (SE) | (SE) |
|---|---|---|---|---|
| Autonomy | Rich | (0.032) | (0.035) | 0.046 (0.015) |
| Autonomy | Survey | (0.036) | (0.038) | 0.044 (0.015) |
| People–Things | Rich | 0.131 (0.017) | 0.055 (0.032) | 0.075 (0.025) |
| People–Things | Survey | 0.119 (0.019) | 0.088 (0.030) | 0.031 (0.022) |
| Creative–Routine | Rich | (0.030) | 0.060 (0.038) | (0.027) |
| Creative–Routine | Survey | (0.032) | 0.004 (0.038) | (0.023) |
| Technology | Rich | 0.082 (0.027) | 0.044 (0.032) | 0.038 (0.026) |
| Technology | Survey | 0.117 (0.029) | 0.076 (0.032) | 0.041 (0.019) |
| Model fit | ||||
| Controls-only | Rich | 0.0740 | ||
| Controls-only | Survey | 0.0740 | ||
| Full | Rich | 0.0857 | ||
| Full | Survey | 0.0846 | ||
| Rich | 0.0117 | |||
| Survey | 0.0106 | |||
| Both | 6,896 | |||
| Occupation clusters, | Both | 309 | ||
- •
Note: is obtained from the deviation-contextual model, from the raw-contextual model, and from either parameterization. Minor discrepancies in reflect rounding. AI measures are constructed using . Standard errors are clustered by occupation.
6.3.1 People–Things
People–Things remains the strongest and most stable group-level predictor in the joint analysis. The estimated group effects are 0.131 under the rich prompt and 0.119 under the survey prompt, with comparatively small standard errors. These estimates are similar across prompting strategies and remain clearly positive after adjustment for Autonomy, Creative–Routine, and Technology.
The individual effect is also positive in both models. It is clearly detected under the rich prompt, with an estimate of 0.075 (SE 0.025), but is less precisely estimated under the survey prompt, at 0.031 (SE 0.022). The contextual estimates are positive in both models, but their evidential strength differs: the rich-prompt estimate of 0.055 (SE 0.032) is imprecise, whereas the survey-prompt estimate of 0.088 (SE 0.030) is more clearly separated from zero.
The conclusion that survives most clearly across prompts and across the single and joint analyses is therefore the positive group effect of People–Things. The positive individual association is also present in both prompt versions, although it is more precisely detected under the rich prompt. Evidence for an additional contextual effect is less robust because it is stronger under the survey prompt than under the rich prompt.
Relative to the single-dimension analysis, the joint model does not attenuate the People–Things group association. The single-dimension estimates were 0.100 and 0.104, whereas the joint estimates are 0.131 and 0.119. Thus, People–Things continues to distinguish occupations with higher versus lower job satisfaction even after variation shared with the other three dimensions is accounted for.
6.3.2 Technology
Technology also exhibits positive group effects in both joint models. The estimates are 0.082 (SE 0.027) under the rich prompt and 0.117 (SE 0.029) under the survey prompt. Although the magnitudes differ, the direction and general conclusion are consistent: occupations with higher mean Technology scores tend to have higher mean job satisfaction after adjustment for the other latent dimensions.
The individual Technology coefficients are positive in both models, but the evidence is stronger under the survey prompt. The survey estimate is 0.041 (SE 0.019), whereas the rich estimate is 0.038 (SE 0.026). The contextual coefficients are also positive but less consistently detected, with estimates of 0.044 (SE 0.032) and 0.076 (SE 0.032).
The joint analysis therefore strengthens the evidence for a Technology group effect. In the single-dimension models, the Technology result was prompt-dependent: the survey measure produced a clear positive group estimate, while the rich measure produced a smaller and less precise estimate. In the joint analysis, both group estimates are positive relative to their standard errors. The survey-prompt individual effect also persists. Evidence for a distinct contextual effect remains less certain because it is again more visible under the survey prompt than under the rich prompt.
6.3.3 Autonomy
Autonomy changes substantially between the single and joint analyses. In the repaired single-dimension models, Autonomy displayed positive group effects of 0.062 and 0.054 and positive individual effects of 0.047 and 0.054. The contextual coefficients were close to zero.
After the remaining latent dimensions are included jointly, the positive individual association remains almost unchanged: under the rich prompt and under the survey prompt. In contrast, the group coefficients decline to and , neither of which provides evidence of a remaining independent group association.
Because the individual coefficients remain positive while the group coefficients approach zero, the implied contextual coefficients become negative:
| (17) |
The estimated contextual coefficients are and , but they are not sufficiently consistent or precise to support a strong substantive claim of a negative contextual effect. Rather, the most defensible interpretation is that the positive group association in the single-dimension Autonomy model was shared with the other latent occupational dimensions. Once that shared variation is controlled, the positive respondent-level Autonomy association remains, but no independent positive group effect is evident.
Thus, the stable Autonomy finding is not a positive group effect. It is the positive individual effect. The repair of the measure remains important because it recovered this association consistently under both prompts. The joint model then clarifies that its group-level counterpart is not independent of People–Things, Creative–Routine, and Technology.
6.3.4 Creative–Routine
Creative–Routine remains the least stable dimension. Its joint group coefficients are (SE 0.030) and (SE 0.032), providing no evidence of an independent group association under either prompt. The rich-prompt individual coefficient is negative, (SE 0.027), while the survey-prompt individual coefficient is much smaller, (SE 0.023). The implied contextual coefficients are positive for Rich, 0.060 (SE 0.038), and approximately zero for Survey, 0.004 (SE 0.038).
These results do not support a stable positive association between Creative–Routine and job satisfaction. They also do not support a conclusive contextual effect. The positive rich-prompt contextual coefficient is produced by combining a near-zero group coefficient with a relatively large negative individual coefficient. The survey-prompt measure does not reproduce that decomposition. Accordingly, the rich result should not be interpreted in isolation as evidence that occupational creativity has a positive contextual effect.
The appropriate conclusion is not that Creative–Routine necessarily failed as a measure. Its occupation rankings exhibited strong face validity, and the rich and survey measures identified recognizable contrasts between creative and routine occupations. Rather, the results indicate that the unique portion of Creative–Routine has no stable positive association with job satisfaction after related occupational characteristics are controlled. The conflicting , , and estimates likely reflect a combination of substantial cross-dimensional overlap and limited independent Creative–Routine variation.
The correlation structure among the four latent dimensions helps explain this instability. Under the rich prompt, Creative–Routine correlates 0.769 with Autonomy, 0.435 with People–Things, and 0.502 with Technology. Under the survey prompt, the corresponding correlations are 0.667, 0.514, and 0.493. Creative–Routine therefore shares substantial variation with every other latent dimension, especially Autonomy. The correlations between the corresponding occupation means are also high.
This structure explains why the single- and joint-dimension analyses answer meaningfully different questions. The single-dimension Creative–Routine model captures both its distinctive component and the variation it shares with Autonomy, People–Things, and Technology. The joint model isolates the narrower component of Creative–Routine that is orthogonal to the other three measures. Because the shared component is large, the coefficient can change substantially when the other dimensions enter the model.
6.3.5 Consistency across prompting strategies
Table 9 compares the joint results obtained from the rich and survey prompting strategies. Because the deviation-contextual and raw-contextual models are algebraically equivalent representations of the fitted model, their predicted satisfaction values are identical. Their coefficient-agreement statistics differ, however, because one representation compares , while the other compares .
| Similarity measure | ||
|---|---|---|
| Correlation of eight focal coefficients | 0.864 | 0.721 |
| RMSE of eight focal coefficients | 0.033 | 0.039 |
| Mean absolute difference of eight focal coefficients | 0.025 | 0.032 |
| Correlation of latent-component predicted satisfaction | 0.823 | |
| RMSE of latent-component predicted satisfaction | 0.078 | |
| Correlation of full fitted satisfaction | 0.967 | |
| RMSE of full fitted satisfaction | 0.075 | |
- •
Note: The column summarizes agreement between rich- and survey-prompt coefficients from the deviation-contextual models. The column summarizes agreement from the raw-contextual models. Prediction statistics are identical across the two parameterizations.
Agreement is stronger for the representation than for the representation. The correlation between the eight group and individual coefficients is 0.864, compared with 0.721 for the contextual and individual coefficients. The RMSE and mean absolute difference are also lower for the group-plus-individual representation.
This difference is consistent with the substantive results. The gross group effects are generally larger and more stable, while contextual effects are obtained after separating the group association from the individual association. Their estimation is consequently more sensitive to differences between the rich- and survey-prompt measures. The contextual decomposition should therefore be interpreted with greater caution than the group decomposition.
Despite coefficient-level differences, the two prompting strategies produce similar overall predictions. The latent-component predicted satisfaction values correlate 0.823, and the full fitted values correlate 0.967. Thus, prompt variation affects the allocation of the fitted association among particular latent dimensions and especially among contextual components more than it affects the overall fitted outcome.
6.3.6 Findings that remain consistent across single and joint models
Two conclusions survive both prompt designs and both the single- and joint-dimension analyses.
First, People–Things has the most robust positive group association. Its group effect is positive under both prompts and in both the single models and the joint models.
Second, the repaired Autonomy measure consistently detects a positive individual association, across prompt versions and individual versus joint models.
These effects are quite reasonable to understand: occupational differences in People–things matter to job satisfaction. Individual (within-occupational) differences in Autonomy also influences job satisfaction, given two jobs with the same occupational mean Autonomy.
7 Discussion
Our results support three conclusions about AICOME for contextual analysis.
First, respondent-level AI measures can recover observed survey dimensions with substantial accuracy when rich respondent and job information are available and when the regression signals are strong enough. In single dimensional analysis across the four CFPS validation dimensions, the AI-derived measures generally reproduce explanatory power, coefficient directions, and conclusions on both the group effects and individual effects obtained using the observed survey responses. The contextual effect is usually harder to detect consistently across prompts, but it can also be detected if it is strong enough, such as in the case of weekly hours, where both prompting strategies recover the large negative contextual effect observed in the CFPS data.
Second, the usefulness of AICOME depends critically on the information available to the model. When the information available to the AI is restricted to occupation and basic demographic characteristics, validation performance deteriorates sharply. By contrast, when richer information is available, including employer characteristics, organizational structure, work schedules, tenure, promotion history, and compensation, AI-derived measures closely track the observed survey variables. This boundary condition is substantively important: the method uses information already present in the dataset to infer omitted constructs, but it does not create information from nothing.
Third, the joint-missingness validation identifies an additional boundary condition. In many realistic applications, researchers seek to recover one or a small number of concepts that were not directly measured, while numerous other respondent-level characteristics remain available. The validation results suggest that AICOME can perform well in this setting. However, performance declines when several related dimensions are simultaneously unobserved and therefore cannot be used to predict one another. Under this substantially more demanding scenario, the AI-derived measures recover broad qualitative features of the observed contextual model, but explanatory power is reduced and several coefficient estimates are attenuated or unstable.
The latent-dimension application illustrates the potential payoff but also the limits of direct validation. The single-dimension latent contextual models show total associations between each latent job dimension and job satisfaction, and do not necessarily agree with the results from the joint analysis when several latent directions are simultaneously used in the model. Across different prompt versions and in both single and joint analyses we conducted, People-things are found to have a positive occupation-level effect on job satisfaction, while Autonomy (after a revised construction that prevents AI from misunderstanding) has a positive individual-level (within occupation) effect.
Several limitations follow from this framing. First, respondent-level AI measures are not substitutes for direct survey measurement. They are most defensible when researchers need to recover a limited number of theoretically important concepts from rich existing respondent information. In addition, our experience with the AI construction of the Autonomy dimension leads us to conclude that so far AI can still miss subtleties in the available respondent information, and cannot yet totally replace human judgment, diagnosis, and correction in forming satisfactory measurements. Second, AI-derived measures should not be treated as ground truth. The validation exercises compare substantive inferences from AI-derived and observed survey measures, but the observed survey measures themselves may also be imperfect indicators of the underlying job concepts. Third, favorable agreement in coefficient signs, coefficient magnitudes, contextual decompositions, or RMSE is not mechanically guaranteed by the framework. Such agreement may suggest that the AI and survey indicators share substantial concept-related signal, but is not a premise of the method. Finally, the present analysis evaluates two prompting strategies but does not exhaust all possible model, prompt, or time-based variation. Future work should examine stability across model versions and prompting designs.
Acknowledgment
Wenxin Jiang is partly supported by a discretionary fund from Northwestern University. He thanks School of Social and Behavioral Sciences of Nanjing University for hosting his summer visits. Yuxiao Wu was partly supported by the Major Project of the National Social Science Fund of China (grant number 22&ZD188). The data are from China Family Panel Studies (CFPS), funded by 985 Program of Peking University and carried out by the Institute of Social Science Survey of Peking University.
Generative AI Disclosure
Large language models were used as part of the research methodology to generate respondent-level contextual measures from survey and occupational information. The AI-generated measures analyzed in this study were produced through a predefined prompting and scoring procedure described in the paper. Xuyang Wang used LLM (GPT-5.6 Sol) during the research. Wenxin Jiang also used generative AI tools, primarily Microsoft Copilot, during the research and manuscript preparation process for tasks including coding assistance, drafting and editing text, brainstorming ideas, and checking references. All AI-generated outputs were reviewed, verified, and revised by these authors as necessary, who take full responsibility for the accuracy, interpretation, and content of the manuscript as a result of AI usage.
References
- Acemoglu and Autor (2011) Acemoglu, Daron, and David Autor. 2011. “Skills, Tasks and Technologies: Implications for Employment and Earnings.” In Handbook of Labor Economics, Vol. 4B, edited by Orley Ashenfelter and David Card, 1043–1171. Amsterdam: Elsevier.
- Adcock and Collier (2001) Adcock, Robert, and David Collier. 2001. “Measurement Validity: A Shared Standard for Qualitative and Quantitative Research.” American Political Science Review 95(3):529–546.
- Argyle et al. (2023) Argyle, Lisa P., Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023. “Out of One, Many: Using Language Models to Simulate Human Samples.” Political Analysis 31(3):337–351.
- Autor et al. (2003) Autor, David H., Frank Levy, and Richard J. Murnane. 2003. “The Skill Content of Recent Technological Change: An Empirical Exploration.” Quarterly Journal of Economics 118(4):1279–1333.
- Bisbee et al. (2024) Bisbee, James, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M. Larson. 2024. “Synthetic Replacements for Human Survey Data? The Perils of Large Language Models.” Political Analysis 32(4):401–416.
- Blalock (1984) Blalock, Hubert M. 1984. “Contextual-Effects Models: Theoretical and Methodological Issues.” Annual Review of Sociology 10:353–372.
- Brynjolfsson et al. (2026) Brynjolfsson, Erik, José Ramón Enríquez, Sophia Kazinnik, and David Nguyen. 2026. “Augmenting Survey Data with Generative AI: An Application to Economic Research.” Stanford Digital Economy Lab Working Paper.
- Cronbach and Meehl (1955) Cronbach, Lee J., and Paul E. Meehl. 1955. “Construct Validity in Psychological Tests.” Psychological Bulletin 52(4):281–302.
- Deming (2017) Deming, David J. 2017. “The Growing Importance of Social Skills in the Labor Market.” Quarterly Journal of Economics 132(4):1593–1640.
- Escanciano and Pérez-Izquierdo (2023) Escanciano, Juan Carlos, and Telmo Pérez-Izquierdo. 2023. “Automatic Locally Robust GMM with Machine-Learning-Generated Regressors.” arXiv:2301.10643.
- Eloundou et al. (2024) Eloundou, Tyna, Sam Manning, Pamela Mishkin, and Daniel Rock. 2024. “GPTs are GPTs: Labor Market Impact Potential of LLMs.” Science 384(6702):1306–1308.
- Enders and Tofighi (2007) Enders, Craig K., and Davood Tofighi. 2007. “Centering Predictor Variables in Cross-Sectional Multilevel Models: A New Look at an Old Issue.” Psychological Methods 12(2):121–138.
- Felten et al. (2021) Felten, Edward W., Manav Raj, and Robert Seamans. 2021. “Occupational, Industry, and Geographic Exposure to Artificial Intelligence: A Novel Dataset and Its Potential Uses.” Strategic Management Journal 42(12):2195–2217.
- Holtdirk et al. (2026) Holtdirk, Tobias, Georg Ahnert, Joseph W. Sakshaug, and Anna-Carolina Haensch. 2026. “In-Context Learning for the Imputation of Public Opinion Data with Large Language Models.” arXiv:2606.09351.
- Horton et al. (2023) Horton, John J., Apostolos Filippas, and Benjamin S. Manning. 2023. “Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?” NBER Working Paper No. 31122.
- Ji et al. (2024) Ji, Junyung, Jiwoo Kim, and Younghoon Kim. 2024. “Predicting Missing Values in Survey Data Using Prompt Engineering for Addressing Item Non-Response.” Future Internet 16(10):351.
- Kim and Lee (2023) Kim, Junsol, and Byungkyu Lee. 2023. “AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction.” arXiv:2305.09620.
- Ludwig et al. (2026) Ludwig, Jens, Sendhil Mullainathan, and Ashesh Rambachan. 2026. “Large Language Models: An Applied Econometric Framework.” Annual Review of Economics 18:283–316.
- Messick (1989) Messick, Samuel. 1989. “Validity.” In Educational Measurement, 3rd ed., edited by Robert L. Linn, 13–103. New York: American Council on Education and Macmillan.
- Murphy and Topel (1985) Murphy, Kevin M., and Robert H. Topel. 1985. “Estimation and Inference in Two-Step Econometric Models.” Journal of Business & Economic Statistics 3(4):370–379.
- Pagan (1984) Pagan, Adrian. 1984. “Econometric Issues in the Analysis of Regressions with Generated Regressors.” International Economic Review 25(1):221–247.
- Peterson et al. (2001) Peterson, Norman G., Michael D. Mumford, Walter C. Borman, P. Richard Jeanneret, Edwin A. Fleishman, Kerry Y. Levin, Michael A. Campion, Melinda S. Mayfield, Frederick P. Morgeson, Kenneth Pearlman, Marilyn K. Gowing, Anita R. Lancaster, Marilyn B. Silver, and Donna M. Dye. 2001. “Understanding Work Using the Occupational Information Network (O*NET): Implications for Practice and Research.” Personnel Psychology 54(2):451–492.
- Zhao et al. (2025) Zhao, Jianpeng, Chenyu Yuan, Weiming Luo, Haoling Xie, Guangwei Zhang, Steven Jige Quan, Zixuan Yuan, Pengyang Wang, and Denghui Zhang. 2025. “Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation.” arXiv:2509.06337.
Appendix
Appendix A Variables and Analytic Roles
This appendix documents the variables used throughout the empirical analyses. Let denote job satisfaction, denote occupation membership, denote the regression control vector, and denote the four observed CFPS measures used for validation. The regression control vector is conceptually distinct from the AI feature sets described in Appendix B. The controls enter only the regression analyses, whereas the feature sets determine the information supplied to the language model when generating respondent-level AI measures.
Throughout the reported analyses, the default control vector was
For analyses involving weekly hours as the focal validated construct, weekly hours was excluded from the control vector to avoid controlling for the construct being evaluated. For other analyses (e.g., for the latent dimensions: autonomy, people-things, creative-routine, and technology), weekly hours was included as an additional control because it was not itself a focal construct.
Occupation Titles.
Occupation titles were supplied to the language model whenever available and served as the primary contextual signal in AI measurement. For all labelled CFPS variables, including occupation codes, we extracted the human-readable category labels from the metadata attached to the survey variables. Occupation codes (qg303code) were therefore represented in the prompts by their corresponding occupation titles rather than by numeric codes.
Appendix B AI Feature Sets and Prompt Protocols
This appendix documents the feature sets, exclusion rules, and prompting procedures used to construct respondent-level AI measures.
B.1 Target Concepts
The framework considers eight job-related concepts:
The first four concepts have corresponding CFPS survey measures and are used for validation. The remaining four concepts do not have direct validation counterparts and are treated as latent dimensions.
B.2 Prompt Styles
For each concept , two AI measures were generated:
The rich-prompt version asked the model to evaluate a broader conceptual construct. For example, the rich computer-use prompt asked how much a person’s job requires computers, digital tools, office software, information systems, data processing, programming, or computer-based communication. The rich latent-dimension prompts similarly used construct-oriented descriptions emphasizing autonomy, people orientation, creativity, and technology.
The survey-prompt version instead asked the model to predict how the respondent would answer a survey-style question. For example, the survey computer-use prompt asked whether the respondent uses a computer for the current job, while the survey foreign-language prompt asked whether the respondent uses a foreign language for the current job.
Agreement between the rich and survey versions provides evidence regarding robustness of the resulting AI measures.
B.3 Response Scales
Computer use, foreign-language use, management responsibility, autonomy, people orientation, creative work, and technology intensity were measured on five-point scales.
Weekly hours was treated differently. Rather than generating a five-point rating, the AI model was asked to estimate the respondent’s average weekly work hours. Consequently, weekly hours entered subsequent analyses as a continuous variable.
B.4 Feature Sets and Exclusion Rules
For each concept , let
denote the concept-specific feature set and
denote the corresponding exclusion set.
The main specification uses
Exclusion rules are concept specific. Variables excluded for one target concept remain available for other concepts unless explicitly excluded there as well.
For the four validated concepts (), a simultaneous-missingness specification was also considered. Define
The AI4 feature sets are then
Thus any variable excluded for computer use, foreign-language use, weekly hours, or management responsibility is excluded from all four AI4 prediction tasks.
Finally, a reduced-information specification was used for sensitivity analysis:
This specification removes essentially all job-specific respondent information beyond occupation membership and basic demographics.
B.5 Concept-Specific Feature Sets
Table 10 summarizes the role of each variable in the concept-specific feature sets.
Entries are coded as:
- •
F: included as a feature in ,
- •
E: excluded due to leakage protection (),
- •
*: unused for that concept.
The AI4 column indicates whether the variable belongs to the unioned exclusion set .
| Comp | Lang | Hours | Mgmt | Auto | People | Creative | Tech | |
|---|---|---|---|---|---|---|---|---|
| Occupation | F | F | F | F | F | F | F | F |
| Age | F | F | F | F | F | F | F | F |
| Gender | * | * | F | F | * | * | * | * |
| Education degree | F | F | F | F | F | F | F | F |
| Urban residence | F | F | F | F | * | * | * | * |
| Marital status | * | * | F | F | * | * | * | * |
| Employer type | F | F | F | F | F | F | F | F |
| Industry | * | * | F | F | * | * | * | * |
| Work location | F | * | * | * | * | * | * | F |
| Province | * | F | * | * | * | * | * | * |
| Organization size | F | F | F | F | F | * | F | F |
| Promotion | * | * | F | F | F | * | F | * |
| Expected promotion | * | * | F | F | * | * | * | * |
| Income | * | * | F | F | * | * | * | * |
| Labor contract | * | * | F | F | * | * | * | * |
| Job tenure | * | * | F | F | * | * | * | * |
| Party membership | * | * | * | F | F | * | * | * |
| Management position | F | F | F | E4 | F | F | F | F |
| Direct reports | F | * | F | E4 | F | F | F | * |
| Number of subordinates | F | * | F | E4 | F | * | * | * |
| Computer use | E1 | * | * | * | * | * | * | F |
| Foreign-language use | * | E2 | * | * | * | * | * | F |
| Weekly hours | * | * | E3 | * | F | * | * | * |
| Time flexibility (Only used in the initial analysis) | * | * | * | * | F | * | F | * |
| Night shift | * | * | F | * | F | * | * | * |
| Weekend work | * | * | F | * | F | * | F | * |
| On-call work | * | * | F | * | F | * | F | * |
Comp = computer use; Lang = foreign-language use; Mgmt = management responsibility; Auto = autonomy; Tech = technology intensity. F = included feature; E = excluded because of leakage protection; * = unused. AI4 uses the union exclusion set .
Appendix C CFPS Validation Measures
The first four concepts, have corresponding observed CFPS validation measures
These measures are generated from survey questions administered in the CFPS occupational module.
C.1 Computer Use
The validation measure is based on question QG19:
Do you use a computer for your current job?
Response categories:
- 1.
Yes
- 2.
No
This question serves as the observed survey measure for the computer-use concept .
C.2 Foreign-Language Use
The validation measure is based on question QG18:
Do you use a foreign language for your current job?
Response categories:
- 1.
Yes
- 2.
No
This question serves as the observed survey measure for the foreign-language-use concept .
C.3 Weekly Hours
The validation measure is based on question QG6:
Excluding lunch break, but including paid or unpaid extra working hours, how many hours per week on average did you work for this job in the past 12 months?
Responses are recorded as average weekly work hours. The questionnaire instructs interviewers to convert minutes to hours and retain one decimal place.
This question serves as the observed survey measure for the weekly-hours concept .
C.4 Management Responsibility
The validation measure is based on question QG14:
Do you have management duty for this job?
The CFPS interviewer instructions define management duty as:
Occupation in the organization that has official management function such as the chief of a section, the director of a department, the head of a bureau, manager, and so on.
Response categories:
- 1.
Yes
- 2.
No
This question serves as the observed survey measure for the management-responsibility concept .
Appendix D Prompt Definitions and Prompt Templates
This appendix documents the exact prompt definitions and prompting templates used to generate respondent-level AI measures.
For each concept , two prompt styles were used:
The rich version asks the language model to evaluate a broader conceptual construct. The survey version asks the model to predict how the respondent would answer a survey-style question.
D.1 Rich Prompt Definitions
Computer Use ()
How much does this person’s job require computers, digital tools, office software, information systems, data processing, programming, or computer-based communication?
Scale: 1 = not at all; 2 = a little; 3 = moderately; 4 = a lot; 5 = very much.
Foreign-Language Use ()
How much does this person’s job require use of a foreign language for reading, writing, speaking, translation, interpretation, international communication, or foreign-language documents?
Scale: 1 = not at all; 2 = a little; 3 = moderately; 4 = a lot; 5 = very much.
Weekly Hours ()
Excluding lunch break, but including paid or unpaid extra working hours, how many hours per week on average did this person work for this job in the past 12 months?
Consider regular work schedules, overtime work, workload, managerial responsibilities, supervisory duties, workload intensity, promotion pressure, schedule demands, and expected time commitment.
Return the number of work hours per week.
Management Responsibility ()
CFPS survey item:
Does this person have management duty for this job?
Management duty refers to positions with official management functions, such as section chief, department director, bureau head, manager, and so on.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
Autonomy ()
How much freedom, discretion, independence, self-direction, and decision-making authority does this person likely have over how work is performed?
Scale: 1 = no freedom; 2 = little freedom; 3 = moderate freedom; 4 = much freedom; 5 = very much freedom.
People Orientation ()
Is this person’s job more people-oriented or things-oriented?
People-oriented work involves teaching, helping, caring, persuading, serving, coordinating, managing, or communicating with people.
Things-oriented work involves machines, tools, equipment, materials, production processes, physical systems, or technical operations.
Scale: 1 = entirely things-oriented; 2 = mostly things-oriented; 3 = mixed; 4 = mostly people-oriented; 5 = entirely people-oriented.
Creative Work ()
Is this person’s job more creative/nonroutine or routine/standardized?
Creative work involves imagination, new ideas, flexible judgment, problem solving, artistic production, or innovation.
Routine work involves standardized, repetitive, predictable tasks following fixed procedures.
Scale: 1 = highly routine; 2 = mostly routine; 3 = mixed; 4 = mostly creative/nonroutine; 5 = highly creative/nonroutine.
Technology Intensity ()
How technologically intensive is this person’s job overall?
Technology-intensive work involves digital technology, engineering, scientific equipment, automation, electronics, information systems, or technical expertise.
Scale: 1 = not technological; 2 = slightly technological; 3 = moderately technological; 4 = highly technological; 5 = extremely technological.
D.2 Survey Prompt Definitions
Computer Use ()
CFPS survey item:
Do you use a computer for your current job?
Based on the available information, predict this person’s likely answer.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
Foreign-Language Use ()
CFPS survey item:
Do you use a foreign language for your current job?
Based on the available information, predict this person’s likely answer.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
Weekly Hours ()
CFPS survey item:
Excluding lunch break, but including paid or unpaid extra working hours, how many hours per week on average did you work for this job in the past 12 months?
Based on the available information, predict this person’s most likely answer.
Return the number of weekly work hours.
Management Responsibility ()
Survey question:
Do you have management duty for this job?
Based on the available information, predict this person’s likely answer.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
Autonomy ()
Survey question:
Does your job give you freedom in how work is done?
Based on the available information, predict this person’s likely answer.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
People Orientation ()
Survey question:
Does your work mainly involve people?
Based on the available information, predict this person’s likely answer.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
Creative Work ()
Survey question:
Does your work require creativity or new ideas?
Based on the available information, predict this person’s likely answer.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
Technology Intensity ()
Survey question:
Does your work require advanced technology?
Based on the available information, predict this person’s likely answer.
Scale: 1 = definitely no; 2 = probably no; 3 = uncertain; 4 = probably yes; 5 = definitely yes.
Appendix E Prompt Templates
All respondent-level AI measures were generated using model gpt-4o-mini and prompt templates that combined:
- 1.
a target concept definition;
- 2.
an occupation title;
- 3.
a concept-specific feature set;
- 4.
instructions governing the response task.
The exact variables supplied to the language model for concept are documented in Appendix B.
Occupation information was always supplied whenever available. If occupation information was unavailable, the model was instructed to rely on the remaining respondent information.
The prompts explicitly instructed the model to evaluate job characteristics rather than job satisfaction, happiness, or mental health.
E.1 Main AI1 Template
The principal measurement design used
For each concept (the latter 4 are latent concepts), the language model received a concept-specific profile block constructed from .
The prompt instructed the model:
You are simulating responses to job-characteristic survey items, not job satisfaction.
Answer as a worker with the occupation and characteristics below.
Do not infer overall satisfaction, happiness, or mental health.
Use the concept-specific profile block for the target dimension.
Occupation is always included when available.
If occupation is not reported, rely on the other respondent characteristics.
For every concept-specific profile block, the prompt contained the instruction:
Pay particular attention to these variables for this dimension.
Return values for the 8 target concepts.
Use the scale specified for each concept.
Most concepts use a 1-5 scale.
weekly_hours must be returned as the number of work hours per week, not as a 1-5 score.
E.2 AI4 Simultaneous-Missingness Template
For the four validated concepts,
an alternative AI4 design was constructed.
The feature sets were
where
The AI4 template was identical to the AI1 template except that computer use, foreign-language use, weekly hours, and management responsibility were treated as simultaneously unavailable.
The prompt instructed the model:
Return values for computer use, foreign-language use, weekly hours, and management responsibility.
Treat all four concepts as unobserved latent quantities.
Variables excluded for any one of the four concepts have been removed from all profile blocks.
Use the remaining respondent information to generate responses.
E.3 Reduced-Information AI4 Template
The reduced-information sensitivity analysis used the same AI4 framework while replacing the concept-specific feature sets with
Only occupation and basic demographic characteristics were supplied to the model.
Job-specific variables, work-history variables, supervisory variables, schedule variables, income measures, and other detailed respondent information were omitted.
The purpose of this specification was to evaluate the extent to which AI measurement depends on access to rich respondent information beyond occupation and demographics.
E.4 Common Processing Rules
Several coded categorical variables were converted into human-readable text before prompt construction. These included schedule flexibility, night-shift frequency, and weekend-work frequency.
Weekly hours was treated as a quantitative construct and was generated directly as a numerical estimate of average weekly work hours.
All other concepts were measured on five-point scales.
The resulting AI measures were subsequently used to construct respondent-level and occupation-level quantities in the analyses reported in the main text.
E.5 Common Cleaning Rules
haven_labelled variables were first converted to ordinary numeric values to do regression. CFPS special missing-value codes were treated as missing. Binary yes/no variables were converted to 1–0 indicators. Nonnegative quantitative variables were restricted to nonnegative values. Invalid occupation codes were removed before constructing occupation-based variables and matching occupation titles.
| Variable | Role | Regression Construction | AI Representation |
|---|---|---|---|
| Job satisfaction qg406 | Five-point satisfaction scale; valid responses retained. | Not used. | |
| Occupation qg303code | , | Invalid occupation codes removed. | Occupation titles obtained from the CFPS value labels. |
| Computer use qg19 | , | Binary indicator (1 = yes, 0 = no). | Human-readable yes/no description. |
| Foreign-language use qg18 | , | Binary indicator (1 = yes, 0 = no). | Human-readable yes/no description. |
| Weekly hours qg6 | , | Finite nonnegative weekly hours. | Weekly hours information supplied directly. |
| Management duty qg14 | , | Binary indicator (1 = yes, 0 = no). | Human-readable yes/no description. |
| Age age | , | Continuous age. | Age value supplied directly. |
| Gender gender | , | Binary indicator. | Human-readable gender description. |
| Education degree cfps2022edu | Not used. | Educational-attainment description. | |
| Education years cfps2022eduy | Finite nonnegative years. | Not used. | |
| Marital status qea0 | , | Married indicator (1 = married, 0 = other status). | Marital-status description. |
| Hukou qa301 | Agridultural hukou indicator (1 = agricultural hukou, 0 = non-agricultural/residence hukou) | Not used. | |
| Urban residence urban22 | , | Binary urban indicator. | Urban/rural description. |
| Self-rated health qp201 | Five-point scale reversed so larger values indicate better health. | Not used. | |
| Income incomeb | , | For nonnegative income, is used in regressions. | Nonnegative income value is supplied to the language model. |
| Employer type qg2 | , | Public/SOE indicator: 1 = government, public institution, or state-owned enterprise. | Employer-type description. |
| Work location qg20 | , | Outdoor-work indicator. Transportation-based workplaces treated as missing. | Work-location description. |
| Organization size qg16 | Finite nonnegative value. | Organization-size information. | |
| Promotion qg15 | Management, technical, or joint promotion coded as promotion. | Promotion history description. | |
| Expected promotion qg1501 | Retained as categorical information. | Expected-promotion description. | |
| Job tenure qg2032 | Binary indicator. | Job-tenure description. | |
| Labor contract qg5 | Binary indicator. | Labor-contract description. | |
| Direct reports qg17 | Binary indicator. | Direct-reports description. | |
| Number of subordinates qg1701 | Finite nonnegative value. | Number-of-subordinates information. | |
| Party membership qn4001 | Binary indicator. | Party-membership description. | |
| Time flexibility qg604 | Not used in regressions. | Three schedule-flexibility categories converted into human-readable descriptions. This variable is used only in the initial analysis. | |
| Night shift qg601 | Not used in regressions. | Human-readable frequency description. | |
| Weekend work qg602 | Not used in regressions. | Human-readable frequency description. | |
| On-call work qg603 | Not used in regressions. | Human-readable description. | |
| Industry qg302code | Not used in regressions. | Industry description. | |
| Work location province derived from provcd22, qg301, qg301a_code, jobclass | Not used in regressions. | Work location province in text. This variable uses the residential province (provcd22) when it is valid and either (1) qg301 indicates that the work location is within mainland China and no valid alternative province (qg301a_code) differs from provcd22, or (2) jobclass indicates family agricultural work. If a valid alternative work province (qg301a_code) differs from provcd22, qg301a_code takes precedence. |
Appendix F Occupation-Level Comparison Methods
To benchmark the respondent-level AI measures proposed in the main text, we also considered two occupation-level approaches based solely on occupation titles. Unlike the respondent-level measures, these benchmarks do not use respondent-specific information and therefore assign the same score to all individuals within a given occupation. A list of occupation titles originated from the labels attached to the qg303_code variable in the CFPS data is used for these occupation-level approaches. Occupation titles served as the sole information source for both benchmark methods.
F.1 Occupation-Level AI Scores
The first benchmark uses direct AI scoring of occupation titles with model rm gpt-4o-mini. For each occupation, the language model was asked to evaluate the typical characteristics of that occupation using only the occupation title as input. The prompts instructed the model to focus on the occupation itself and not on job satisfaction. No respondent-level characteristics were supplied.
For the seven ordinal concepts (computer use, foreign-language use, management responsibility, autonomy, people orientation, creative work, and technology intensity), the model returned a score on a five-point scale using the same rich concept definitions employed in the respondent-level AI measures. For weekly hours, the model instead estimated the average number of hours worked per week by a typical worker in the occupation. The resulting occupation-level scores were merged back to individuals using occupation codes, so that all workers within the same occupation received identical benchmark scores.
F.2 Occupation Embedding Scores
The second benchmark uses occupation embeddings with model text-embedding-3-large. Each occupation title was converted into a vector representation using an embedding model. To construct these embeddings, the model was asked to generate a neutral representation of the occupation’s typical work content, responsibilities, and work setting, without reference to any particular survey dimension or outcome.
For each target concept, a semantic direction was constructed from positive and negative textual descriptions representing opposite ends of the underlying dimension (e.g., computer-intensive versus non-computer work). Occupation scores were obtained by projecting each occupation embedding onto the corresponding semantic direction. Higher projected values indicate greater semantic similarity to the positive pole of the concept.
As with the occupation-level AI scores, the embedding benchmark depends only on occupation titles and therefore assigns the same score to all workers sharing the same occupation.
F.3 Relation to Respondent-Level AI Measures
Both occupation-level benchmarks utilize occupation titles alone. In contrast, the respondent-level AI measures proposed in the main text combine occupation information with the th respondent’s specific characteristics through the feature sets . Consequently, the proposed method allows workers within the same occupation to receive different predicted scores, whereas the occupation-level benchmarks do not.