How India Plans to Count Caste This Time — And What Went Wrong Earlier
India’s ongoing Census will use an open-ended question to record caste, allowing respondents to state their caste rather than choosing from a predetermined list. The approach has raised concerns about duplicate names and an unwieldy dataset, similar to problems encountered during the 2011 Socio-Economic and Caste Census.
Written by
Jyoti Mukherjee

New Delhi: India's ongoing Census is set to attempt caste enumeration in a way that could significantly influence how caste data is collected and interpreted in the country.
Instead of providing respondents with a predetermined list of castes or a drop-down menu, the Registrar General of India has decided to use an open-ended question on caste.
Under this approach, respondents will be asked to state their caste, which will then be recorded by the Census enumerator.
The method is straightforward in principle, but it comes with a major administrative challenge: the same caste may be identified by different names, spellings or surnames.
This could result in a very large number of separate entries, making it difficult to determine which entries actually represent the same caste or social group.
A similar problem emerged during the 2011 Socio-Economic and Caste Census (SECC), when the open-ended approach resulted in an enormous number of distinct caste names.
Why the open-ended approach could create problems
Under the current system, there will be no pre-given list of castes for respondents to select from.
If a person identifies their caste through their surname, the enumerator is expected to record the answer as stated.
That creates the possibility of different spellings or variations becoming separate entries.
A senior IAS officer cited in the report explained that if someone identifies their caste by their surname, the Census official would be required to record it as the respondent's caste. Even a misspelling could potentially become another entry in the caste column.
The difficulty is compounded by the fact that surnames do not necessarily correspond to a single caste.
For example, names such as Rajput, Thakur, Singh and Kshatriya may be used interchangeably by some respondents.
At the same time, the same word can have different meanings in different regions.
The surname Singh, for instance, is used across multiple castes and categories.
Similarly, the surname Verma can refer to people belonging to different communities, including Kayasths, Kurmis, Jats and Sonars.
The report also points out that Rajput in the Bundelkhand region can refer to the Other Backward Class Lodh community.
These variations make caste enumeration considerably more complicated than simply asking people to state their identity.
SC and ST enumeration follows a different model
The current caste-counting exercise differs from the way Scheduled Castes and Scheduled Tribes are enumerated.
SC and ST lists are already officially recognised and comprehensive lists are used during Census exercises.
The Ministry of Social Justice and Empowerment lists 1,208 Scheduled Castes, with the recognised communities varying across states.
A 2022 Press Information Bureau year-end release listed 730 Scheduled Tribes in India.
Because these categories have recognised lists, the data collected for SCs and STs can be organised more systematically.
The absence of a similar predetermined list for the broader caste enumeration creates a different set of challenges.
India has faced this problem before
The difficulty of accurately counting caste is not new.
The colonial government also encountered serious problems when caste was enumerated during the 1931 Census, the last Census in which comprehensive caste data was published.
The 1931 Census demonstrated how subjective caste identity could become during enumeration.
Some communities consolidated themselves into larger groups, while others adopted new names or identities.
In some cases, communities sought to change how they were officially classified in an attempt to improve their social standing or strengthen their numerical representation.
The Aad Dharmis example
One example from the 1931 Census involved sections of leather-working communities among Dalits in Punjab.
They adopted the identity of Aad Dharmis, referring to themselves as followers of an original or pre-Aryan religion of India.
The category had not existed in the 1921 Census.
By 1931, around 418,789 people identified themselves as Aad Dharmis.
The number was roughly comparable to the Christian population of Punjab at the time.
Aad Dharmis accounted for approximately 1.5% of Punjab's population and around one-tenth of the population classified as low castes in the province.
The Census report also noted that literacy among Aad Dharmis was more than twice that among other members of the low-caste population.
The same broader trend of communities adopting new collective identities appeared elsewhere.
Names such as Adi Dravida, Adi Andhras and Adi Karnatakas emerged in different parts of India.
Communities also tried to consolidate their numbers
The 1931 Census recorded another important phenomenon: multiple castes attempting to consolidate themselves under a single identity.
The objective could be to increase numerical strength or claim a higher social status.
One prominent example involved communities associated with pastoral and cattle-rearing occupations.
The Census report noted efforts to bring groups such as Ahirs, Goalas, Gopis and Idaiyans under the broader Yadava identity.
The report said this consolidation movement had already been effective to some extent by 1921.
Similar efforts were reported among occupational groups such as carpenters, smiths and goldsmiths.
Some sought to use a common identity such as Vishwakarma or Jangida, sometimes adding descriptions intended to associate the group with higher social status.
The Census also recorded cases where groups sought to identify themselves as Brahmins or Rajputs.
In some cases, the same community sought one identity in one province and another identity elsewhere.
What happened in 1941 and 1951?
The comprehensive caste-counting approach did not continue unchanged.
In the 1941 Census, caste information was collected, but caste was not included in the final tabulation.
Census Commissioner M W M Yeatts argued that the enormous cost of producing detailed caste tables was no longer justified, particularly during the financial constraints associated with the Second World War.
After Independence, the approach changed further.
In the 1951 Census, the government led by Prime Minister Jawaharlal Nehru decided not to enumerate caste, reflecting the new country's emphasis on equality and secularism.
As a result, the 1931 Census remained the last comprehensive caste census before the issue returned to national policy debates decades later.
The 2011 caste census problem
The 2011 Socio-Economic and Caste Census provides perhaps the most direct warning for the current exercise.
The UPA government conducted the caste enumeration using an open-ended method.
The result was a massive number of entries.
The final dataset contained approximately 46.7 lakh distinct caste names.
That figure was dramatically higher than the 4,147 castes recorded in the last comprehensive caste census of 1931.
The scale of the problem meant that the government ultimately withheld the raw caste data.
The experience demonstrated the difficulty of collecting caste information without a standardised classification system.
Bihar's 2023 caste survey
The experience of Bihar provides another example.
When Bihar conducted its statewide caste survey in 2023, officials reportedly prepared a caste list before conducting the exercise.
An IAS officer familiar with the process told The Indian Express that having a list was important from an administrative perspective.
Without a standardised list, similar caste identities could appear as separate categories.
For example, respondents identifying themselves as Bhumihar and others identifying themselves as Bhumihar Brahmins could potentially be counted as two distinct categories.
The Bihar survey therefore attempted to address some of the classification problems associated with open-ended enumeration.
However, the Bihar exercise also faced questions about how communities were grouped.
Former Bihar Deputy Chief Minister Sushil Modi had raised concerns over the methodology used to prepare the caste lists.
He pointed to examples where several Yadav sub-castes were consolidated under one heading while some Kushwaha sub-castes, including Dangis, were counted separately.
Similar questions were raised regarding the treatment of sub-castes among fishing communities.
This highlights another difficult question: even when a predetermined list is used, deciding which communities should be combined and which should remain separate can itself become politically and socially sensitive.
What could be the solution?
Some experts believe that the Census should use a predetermined list rather than leaving caste identification entirely open-ended.
Former Indian Council of Social Science Research chairman S K Thorat, who was part of an expert committee on caste data in Telangana, argued that a comprehensive list would help reduce classification errors.
He pointed out that India already has recognised SC, ST and OBC lists.
According to him, a similar list could be prepared for castes in the general category, creating a more exhaustive classification framework.
Thorat acknowledged that even such a list would not eliminate all errors.
He estimated that there could still be a small margin of error of around 2-3%.
However, he argued that this would be preferable to an entirely open-ended system that could produce millions of separate entries.
What about people who do not identify with caste?
A comprehensive caste questionnaire would also need to account for people who do not identify with a caste.
Thorat suggested that the Census could include options such as "no caste" and "no religion".
He also proposed separate questionnaires designed to capture the specific forms of exclusion experienced by different communities.
For Scheduled Castes, such questions could address forms of exclusion historically associated with untouchability.
For Scheduled Tribes, the questionnaire could address issues connected to physical and geographical isolation.
He argued that the questions should also capture hierarchies and inequalities within SC and ST communities because some groups may remain significantly more deprived than others.
Why the method matters
The way caste is counted is not merely a technical issue.
The resulting data could influence how policymakers understand India's social structure and inequalities.
But if similar caste groups are counted separately because they use different names, or if different communities are merged under a single label without a clear methodology, the final numbers could be difficult to interpret.
The experience of 2011 demonstrates the scale of this problem.
The difference between 46.7 lakh reported caste names and the 4,147 castes recorded in 1931 shows how dramatically the number of entries can expand when respondents are allowed to provide their own descriptions.
At the same time, a rigid list also has limitations because caste identities can vary by region and communities may use different names for themselves.
The challenge is therefore to create a system that is detailed enough to capture India's social diversity while standardised enough to produce usable data.
A difficult balancing act
India's current caste enumeration is attempting to navigate a problem that has existed for more than a century.
The 1931 Census showed how communities could adopt new identities, merge with other groups or seek higher-status classifications.
The 2011 SECC demonstrated how an open-ended approach could produce an enormous number of separate entries.
The 2023 Bihar survey showed that using lists can reduce some administrative difficulties but can also raise questions about how sub-castes are grouped.
The ongoing Census therefore faces a delicate balancing act.
A completely open-ended approach may capture how people identify themselves but create major difficulties in standardising the data.
A predetermined list may make the data easier to analyse but risks overlooking regional variations, alternative names and emerging identities.
The eventual usefulness of India's caste data will depend not only on how many people are counted, but also on how consistently their responses can be classified and interpreted.
Keep reading
More in National

National
Delhi Slum Rehab Plan 2047: High-Rises, More Saleable Space to Attract Developers
Delhi’s draft Master Plan 2047 proposes a major change to the slum rehabilitation model by allowing higher-density construction and a larger…
National
Kolkata Airport Fire: Blaze Breaks Out at Old ATC Building, Flight Operations Unaffected
A fire broke out at the old Air Traffic Control building of Kolkata’s Netaji Subhas Chandra Bose International Airport on Tuesday. The blaze…

National
Haldia Hit by Heavy Rain: 134 mm Rainfall Leaves Residents Facing Waterlogging
Haldia recorded 134 mm of rainfall in 24 hours amid intense monsoon activity across South Bengal. Heavy rain has raised concerns over waterl…

National
Manasa Puja 2026: Government Holiday Announced in Purulia and Bankura
The West Bengal government has announced a public holiday in Purulia and Bankura on the occasion of Manasa Puja. The announcement was made b…
National
Cauvery Wildlife Sanctuary Firing: Forest Officials Booked After Three Men Killed
Protests have erupted in Karnataka’s Chamarajanagar district after three men were killed in a firing incident involving forest personnel at…

National
8 Killed in West Bengal Hotel Fire as Locked Emergency Exit Traps Guests
Eight people, mainly pilgrims, were killed and several others injured after a devastating early-morning fire broke out at Hotel Bideshini ne…
