Abstract
Public transgender history in mainland China increasingly reaches readers through search, platform indexes, digital archives, and automated summaries. The first representation a reader encounters is often a compressed layer rather than the underlying record: a title, date, source label, search snippet, platform tag, machine-produced description, or small set of structured fields. Compression makes discovery efficient, while also deciding which historical facts receive attention first. The 2009 national rule for sex-reassignment surgery can appear as an official governance document with a document number and publication date, or as a news result centered on age, marital status, and required certificates. The 2017 national transgender survey can appear as a complete research project or as a memorable statistic detached from its recruitment frame. The Mr C employment case can enter discovery through labor arbitration, unlawful termination, personality rights, or equal-employment rights. Peking University Third Hospital (PUH3) can surface as a current clinic guide, a research announcement, a historical team record, or an archive page with an automated summary. Jin Xing’s 2011 judging controversy and her 2015 talk-show role generate sharply different public-history frames. A platform copy of A Day of Trans can foreground uploader metadata while the creator-controlled page foregrounds authorship and production.
This article conducts a bounded provenance audit of six fixed historical clusters using publicly accessible pages, public-web discovery results, archive records, and platform pages observed on September 8, 2026. The unit of analysis is the retrieval representation, rather than an individual user. For each object, the audit records historical identity, source role, page metadata, query frame, surfaced title or summary, machine-generated status, and the historical inference a reader could form before opening the source. The article proposes two reusable tools. The first is a discovery compression chain: historical object → source record → page metadata → index representation → query frame → result snippet/automated summary → reader inference. The second is a retrieval context ledger that keeps object_identity, source_role, historical_date, institutional_actor, scope_or_denominator, identity_provenance, legal_stage, and summary_status in separate fields. Automated summaries receive an additional summary provenance contract recording generation status, source object, source role, observation time, known uncertainty, original-link status, and archival relation.
Across the fixed sample, the recurring historical risk is often context decoupling rather than simple disappearance. A date may survive while the institutional stage fades. A percentage may survive while its denominator and sampling frame fade. A person’s name may survive while the provenance of an identity label fades. A video may remain available while creator provenance and uploader provenance trade places in the interface. The same sample also reveals constructive designs. Archive pages that visibly label machine-generated summaries, expose unknown fields, preserve upstream links, and show archive dates create a transparent discovery layer. Vertical search can expose source profiles and verification dates before a user clicks a result. Discoverability therefore has two levels: finding an object, and finding it with enough provenance context to interpret it responsibly.
Keywords: transgender history in mainland China; search; automated summaries; digital archives; platform visibility; provenance; retrieval context; AI-assisted navigation
1. Search results have become a second cover for historical records
Web-history research has long separated link persistence from content persistence. A URL can remain live while the represented content changes, and an original page can disappear while copies, citations, and archive captures continue to carry the record forward (Klein et al. 2014; Jones et al. 2016). Digital-history scholarship extends the point: interfaces, crawl dates, versions, and platform infrastructures belong to the historical conditions through which researchers encounter web sources (Brügger 2012; Fickers 2020; Nix & Decker 2021). These conditions matter greatly for transgender public history in mainland China because many consequential records first circulated as news articles, community reports, hospital pages, advocacy archives, video-platform pages, or institutionally maintained web documents.
Search and automated summarization add another layer above those sources. Readers commonly see a title, host domain, date, snippet, and a small amount of structured metadata before deciding whether to open the underlying page. Platform studies describe visibility as an outcome shaped by platform architecture, metadata, ranking, and user practice (Poell, Nieborg & van Dijck 2019). Research with transfeminine creators shows that visibility also carries opportunities and exposure costs, motivating creators to build folk theories about ranking and recommendation systems (DeVito 2022; Haimson et al. 2021). Everyday algorithm-auditing research demonstrates that repeated observation, controlled comparison, and interface documentation can produce bounded evidence about such systems (Shen et al. 2021; Simpson et al. 2022).
This observatory applies that logic to historical retrieval. The central question is simple: which provenance fields survive into the first public representation of a historical object? Relevant fields include source role, historical date, institutional stage, denominator, provenance of identity language, creator relationship, and machine-generation status. Every result surface has limited space, so compression is an ordinary interface function. The research task is to record what compression keeps in view and what it postpones until a later click.
This question complements GenderLibs’ earlier study of how transgender sources survive web change. That study followed a source → replica → community archive → web archive → discovery index survival stack. The present observatory moves inside the final discovery layer. Once an object is discoverable, search titles, snippets, platform tags, archive metadata, and automated summaries still shape the reader’s first historical frame. Source survival asks whether a route still reaches the object. Retrieval context asks what the reader sees before reaching it.
2. Materials and Method: six fixed clusters and one retrieval context ledger
The observatory uses six clusters that already have multiple public source roles in mainland China’s transgender history: the 2009 Technical Management Specification for Sex Reassignment Surgery (Trial); the 2017 Chinese transgender population survey; the Guiyang Mr C employment case; PUH3’s transgender medicine records; Jin Xing’s television-role history from 2011 through 2015; and the 2021 documentary short A Day of Trans. Each cluster has at least two of the following layers: an original or institutional record, contemporaneous reporting, later platform circulation, a community archive, or a specialist index. That multiplicity makes source-role shifts observable.
The observation date is fixed at September 8, 2026. Materials include public pages from the National Health Commission, hospitals, news organizations, advocacy organizations, digital archives, creator-controlled pages, video platforms, and specialist search services. Each cluster was checked through several semantically related discovery frames, including formal title, person or institution name, event description, topic terms, and historical year. The audit records titles, dates, source roles, and fields that can be re-opened on public pages. Instantaneous rank position is treated as an interface observation tied to this date. Longitudinal rank statistics, personalized-account effects, and geographically localized result differences belong to a later fixed-device, fixed-query, repeated-sampling design.
The retrieval context ledger uses eight core fields:
| Field | Historical question | Common compression effect |
|---|---|---|
object_identity | Which exact document, case stage, program, or film does the result represent? | similarly named objects merge |
source_role | Official release, contemporaneous report, advocacy archive, repost, platform upload, or machine index? | a secondary doorway resembles a source record |
historical_date | When did the event, publication, repost, archive capture, and update occur? | several clocks collapse into one date |
institutional_actor | Who issued, organized, judged, produced, or preserved the record? | institutional relations become generic attribution |
scope_or_denominator | Which sample, subgroup, question, or analytic stage supports a number? | a memorable statistic loses its denominator |
identity_provenance | Does an identity term come from self-description, media, medicine, law, or current editorial description? | historical labels appear to be uniform self-identification |
legal_stage | Arbitration, labor judgment, personality/equal-employment litigation, or later appeal/archive? | several legal tracks become one “victory” |
summary_status | Human-edited text, page text, platform tags, search snippet, or automated summary? | machine language acquires the appearance of source prose |
The design borrows from multi-situated platform research: the same object across several interfaces forms the unit of inquiry, while any one page provides only one part of the evidence (Dieter et al. 2019). Community-archive scholarship adds a second requirement. Description should preserve who is doing the describing, the relationship between archive and community, and the capacity of represented groups to shape their own record (Caswell & Cifor 2016; Caswell et al. 2017; Wakimoto, Bruce & Partridge 2013). The observatory therefore records both what a summary says and what system or source role produced the summary.
3. The 2009 medical rule: a governance document becomes an eligibility checklist
The National Health Commission’s currently accessible record provides a strong governance-oriented representation of the 2009 rule. Its title identifies the document as a Ministry of Health notice issuing the Technical Management Specification for Sex Reassignment Surgery (Trial). The page shows a publication date of November 20, 2009, the document number 卫办医政发〔2009〕185号, the regulatory purpose of reviewing clinical application and safeguarding medical quality and safety, and the November 13 signature date (Ministry of Health 2009). A reader entering through this page sees document type, responsible institution, policy purpose, and two historical clocks.
News-oriented discovery commonly shifts attention toward eligibility: age, marital status, certificates, hospital qualifications, and procedural requirements. That compression serves a practical information need. A reader asking what surgery requirements existed at the time can identify relevant conditions quickly. Historical inquiry requires the governance layer to return: draft discussion, final notice, review target, issuance date, and subsequent regulatory change are separate questions. The source document and an eligibility-centered summary therefore produce different first-order frames around the same historical object.
The Trans Chinese Digital Archive offers a third representation. Its page preserves the upstream National Health Commission URL and lists both archive.md and Internet Archive captures alongside the reproduced text (Trans Chinese Digital Archive 2026a). Project Trans places the rule inside a chronological legal-policy archive (Project Trans 2026). Here, the compact interface foregrounds source URL, archival relation, and legal time. Research on reference rot and content drift explains why this multi-route structure supports long-term recoverability (Klein et al. 2014; Jones et al. 2016). The additional retrieval insight is that one record can sustain several useful compressed covers at once: an eligibility summary, an official governance record, and a provenance-centered archive page.
For this object, the ledger should preserve at least source_role=official_rule, historical_date=2009-11-13/20, institutional_actor=Ministry of Health, and object_identity=卫办医政发〔2009〕185号 plus attached specification. A concise search result can foreground eligibility while those provenance fields remain recoverable on the next layer. Historical note-taking becomes more durable when the researcher saves both the convenient summary and the governance fields.
4. The 2017 national survey: how a memorable number outruns its sampling frame
The 2017 Chinese transgender population survey creates a classic case of numerical compression. The contemporaneous China Development Brief record identifies the survey as a joint project of the Beijing LGBT Center and the Department of Sociology at Peking University and records its public release on November 20, 2017 (Beijing LGBT Center 2017). The survey and later research preserve a larger evidence lineage: 5,677 submitted questionnaires, 2,060 valid responses, and several smaller analytical subsets used in later studies. A single percentage or count travels exceptionally well through search and secondary summaries, while recruitment method, valid-sample rule, identity subgroup, question wording, and analytical filtering require substantially more interface space.
This problem connects directly to research on transgender health-information seeking. Online information can expand access while users still need to evaluate relevance, quality, and applicability (Augustaitis et al. 2021). Community-engaged digital knowledge mobilization likewise depends on both reach and context (MacKinnon, Kia & Lacombe-Duncan 2021). Survey statistics require the same pairing. A result that surfaces one striking number performs discovery work; a responsible research workflow reconnects that number to the underlying denominator and questionnaire frame.
The Trans Chinese Digital Archive’s 2017 survey page is especially useful because it exposes the status of its automation. The page explicitly marks its summary and supplemental information as automatically generated for retrieval and reference, and it provides structured metadata for filename, format, size, MD5, archived date, original link, author, region, date, and tags (Trans Chinese Digital Archive 2026b). Machine-generation status is therefore part of the public evidence. A reader can use the summary to identify the object while retaining a clear cue that exact quantitative claims belong to the PDF or contemporaneous release record.
This design suggests a minimal summary provenance contract. For this page it can record summary_status=machine_generated, source_object=2017 survey PDF, source_role=community_archive, observed_at=2026-09-08, original_url=<recorded source>, archival_relation=archived copy, and uncertainty=<page-declared correction/unknown channel>. Jaillant and Caputo’s work on AI and born-digital archives highlights the importance of interpretive infrastructure around automated access (Jaillant & Caputo 2022). A visible machine-generated label turns that abstract requirement into a concrete user-facing field.
Survey retrieval also benefits from a mandatory scope_or_denominator field. Every count, percentage, risk ratio, or subgroup estimate should point to a sample, identity subset, item, and analytical stage. Search can stay compact while the next layer restores the statistical object from which the number came.
5. The Mr C employment case: one “victory” label contains several legal clocks
The Guiyang Mr C case demonstrates legal-stage compression. A 2016 China News report used period media language in its headline while documenting the labor-arbitration stage, the workplace dispute, and Mr C’s male self-identification in the body (China News 2016). Reporting around the first court outcome emphasized unlawful termination and economic relief while describing a more complicated record around discriminatory causation (RFA 2017a). A later 2017 report focused on a different rights-oriented stage involving employment equality, personality rights, and discrimination (RFA 2017b). Common Language’s later advocacy archive highlighted the significance of “gender identity and gender expression” entering judicial language (Common Language 2020).
A retrieval interface that retains only “first transgender employment discrimination case won” can fuse arbitration, labor-contract adjudication, personality/equal-employment litigation, and later advocacy interpretation into one judicial act. Legal history benefits from a more granular legal_stage chain: workplace event → arbitration → labor judgment → personality/equality-rights judgment → later advocacy archive. GenderLibs’ earlier dual-track study of the case showed that shared workplace facts entered different causes of action and evidentiary tasks. The retrieval layer adds another transformation: search compresses those legal stages back into a clickable sentence.
Identity-language provenance also matters. Historical news headlines contain media terminology characteristic of the period; article bodies may preserve the litigant’s own identity statements; judicial language has another vocabulary; current editorial history often uses a stable explanatory term such as “transgender man.” The ledger can preserve self-identification, contemporaneous-media-label, legal-language, and current-editorial-language as distinct provenance roles. Historical-title fidelity and respectful current description then coexist in the same archive record.
Content-governance research frequently frames platform decisions as trade-offs among visibility, social classification, and contextual interpretation (Jiang et al. 2022; Haimson et al. 2021). Legal-history retrieval has a parallel structure. A shorter label raises recognizability; a fuller procedural chain raises interpretive precision. A source-role ledger lets both functions coexist while keeping the result surface concise.
6. PUH3: current service, institutional history, research news, and archival summary are different source roles
Queries around “PUH3 transgender / 易性症” show a particularly clear source-role shift. PUH3 Department of Plastic Surgery’s current Comprehensive Diagnosis and Treatment Guide is a service-oriented page centered on contemporary clinical access and care pathways (PUH3 Department of Plastic Surgery 2025). A 2026 Department of Endocrinology page is research news, reporting a JAMA Network Open publication and a newer national transgender health survey (PUH3 Department of Endocrinology 2026). A 2018 plastic-surgery department record marks the formation of a multidisciplinary team and functions as an institutional-history node (PUH3 Department of Plastic Surgery 2018). The same institution and overlapping vocabulary therefore lead to at least three source roles: service, research_news, and institutional_history.
The Trans Chinese Digital Archive copy of an older “易性症” treatment guide adds a fourth role: archived_historical_guide. Its automated summary is visibly labeled, while the structured metadata exposes Original Link | [Unknown link(update needed)] and an unknown date (Trans Chinese Digital Archive 2026c). For historical research, the unknown field is valuable evidence. It tells the reader exactly where the archive’s provenance chain currently ends. A fluent summary and an explicit provenance gap can occupy the same page, allowing “text is readable” and “original web provenance is fully reconstructed” to remain separate historical claims.
This case shows why automated-summary quality should be assessed along more than a fluency axis. Archives need traceable fluency: the description helps users identify the object, while metadata shows how far the source chain has been reconstructed. Community-archive scholarship connects description to responsibility toward represented communities (Caswell & Cifor 2016; Caswell et al. 2017). Research on AI and born-digital archives extends that responsibility to machine processing, source selection, and explanatory interfaces (Jaillant & Caputo 2022). The PUH3 archive page demonstrates a practical principle: generated fields and unknown provenance fields should both remain visible.
A PUH3 retrieval ledger therefore needs at least source_role, historical_date, institutional_actor, summary_status, and original_url_status. A 2025 service guide and an older PDF with unresolved original-link provenance can then remain clearly separated even when their titles and clinical vocabulary overlap.
7. Jin Xing: the query frame chooses an identity event or a hosting role
Jin Xing’s television history illustrates another form of compression: the query frame selects which public role becomes the first cover. China News reports from September 2011 organized their headlines around the judging-qualification controversy and explicitly foregrounded her transition history as part of the event frame (China News 2011a, 2011b). A follow-up interview continued to revolve around the reported ban, discrimination, and Jin Xing’s response (China News 2011c). These pages are crucial evidence for studying how television institutions and media language handled a transgender public figure in 2011, because the headline itself forms part of the historical classification system.
By 2012, Sina’s coverage of Jin Xing Hits Mars shifted the central role to “host” (Sina Entertainment 2012). Coverage of The Jin Xing Show in 2015 moved further toward talk-show format, presenting style, and program brand (China News 2015; People’s Daily 2015). The queries Jin Xing 2011 judge and Jin Xing 2015 Jin Xing Show therefore open different role trajectories. The first commonly foregrounds identity controversy and access to a judging position. The second foregrounds hosting authority and a program organized around her name.
Media-visibility scholarship has long distinguished appearance from the structure and quality of representation (Duguay 2016). Research on transfeminine creator visibility adds the role of platform strategy, algorithmic mediation, and community interpretation (DeVito 2022). The Jin Xing case moves this issue into historical retrieval. Researchers need to preserve the query frame if they want to explain why the same person acquires different first-order historical descriptions in the same discovery environment.
A person-history record can therefore add query_frame and role_frame. Examples include query_frame=2011 judging dispute, role_frame=judge/access controversy and query_frame=2015 talk show, role_frame=host/program brand. Taken together, the two records reveal role migration. Each one remains a query-dependent slice.
8. A Day of Trans: work provenance and platform-upload provenance need separate columns
The 2021 documentary short A Day of Trans moves the source-role problem into video platforms. Yennefer Fang’s creator-controlled Vimeo page describes four transgender people across three generations in China, identifies the film’s connection to Transgender Day of Remembrance, and credits Yennefer Fang as writer, director, and producer; it also records a 20-minute runtime and China as the country of production (Fang 2021). This page is close to the work’s creator-controlled provenance.
A 2022 4K upload on Bilibili supplies another kind of historical evidence. The page records an upload date, tags such as culture, people, transgender, and LGBT, and the uploader’s account identity. The uploader’s profile text states that the channel contains documentaries and other films and distinguishes compressed/reposted works from the uploader’s own stance (Bilibili 2022). For circulation history this page is valuable: it records how the film entered a Chinese video-platform environment, which discovery tags accompanied it, and which account supplied that platform copy.
If a retrieval result foregrounds only the Bilibili page, uploader sits visually close to the work and can be mistaken for production provenance. If discovery foregrounds only Vimeo, the later Chinese-platform circulation trail becomes less visible. The clean solution is to keep creator_provenance and platform_upload_provenance in separate fields. Moving-image archival work similarly separates production, premiere, festival circulation, platform circulation, and later curation. Retrieval research adds one more step: a search result’s “source” generally identifies the host page first, while authorship of the underlying work requires its own provenance check.
Social-media data archive research shows how APIs, interface structures, and platform policy determine which metadata future researchers inherit (Acker & Kreisberg 2020). Platformization research similarly connects content visibility to infrastructures and market/platform organization (Poell, Nieborg & van Dijck 2019). For transgender moving-image history, retaining two provenance columns lets platform circulation remain a valuable historical layer while preserving the film’s production relation.
9. A constructive comparison: vertical search can make source role part of the product
Compression can also add context. TSIndex currently presents itself as a specialist search and navigation service for transgender, nonbinary, CDTS, gender-transformation, femboy/nanniang, and related Chinese-language materials. Its public home page reports 39 indexed sources, 39,364 searchable records, and 27,052 restricted records (TSIndex 2026a). Its About page describes AI-assisted navigation, semantic analysis, a knowledge graph, quality and reliability boundaries, and privacy commitments (TSIndex 2026b). These are platform self-descriptions and should be interpreted as such; they also demonstrate a design direction in which source structure becomes visible in the retrieval interface.
The strongest example is TSIndex’s source archive. Each indexed source receives a profile documenting publisher, site function, coverage, suitable uses, source basis, classification, current access state, last verification date, timeliness limits, and significant changes. The page also states that a source profile and existing search records remain preserved when the upstream site becomes temporarily unavailable, times out, returns errors, stops updating, or migrates (TSIndex 2026c). These fields closely resemble the retrieval context ledger. A search product can make “why is this result here?” inspectable rather than presenting only a title and rank.
The design also resonates with participatory-description arguments in community-archive scholarship. Wakimoto and colleagues connect queer archiving with activist and community roles (Wakimoto, Bruce & Partridge 2013), while Caswell and collaborators stress representation and the social capacity to imagine otherwise (Caswell et al. 2017). A vertical retrieval system can therefore be evaluated along two dimensions: how many objects it indexes, and how much provenance context it preserves. The second dimension is particularly important for historical work.
10. Three models: compression chain, context ledger, and summary provenance contract
10.1 Discovery compression chain
The six clusters can be represented with one chain:
historical object → source record → page metadata → index representation → query frame → result snippet/automated summary → reader inference
Information can be reordered at every arrow. A source document foregrounds governance identifiers while a news result foregrounds eligibility. A survey foregrounds methods and sample while a snippet foregrounds one number. A litigation archive foregrounds procedural stages while a search title foregrounds “victory.” A creator page foregrounds production while a video platform foregrounds uploader and tags. Digital-memory scholarship treats networked memory as a continuing process of remediation and rearrangement (Mandolessi 2024). The compression chain provides an operational version of that idea for web retrieval.
10.2 Retrieval context ledger
A minimal ledger for historians, archives, and specialist search systems can record:
1object_id
2query_frame
3observed_at
4result_url
5result_title
6source_role
7historical_date
8institutional_actor
9scope_or_denominator
10identity_provenance
11legal_stage
12creator_provenance
13platform_upload_provenance
14summary_status
15original_url_status
16archive_relation
Every object uses only the fields that matter. Legal cases emphasize legal_stage; surveys emphasize scope_or_denominator; person histories emphasize identity_provenance; video objects emphasize creator and uploader provenance; archive summaries emphasize summary_status and original_url_status. A unified ledger makes empty fields meaningful. A value such as unknown or needs verification becomes an explicit research state with a next step.
10.3 Summary provenance contract
Automated summaries benefit from six additional categories of information:
generated_status: human-edited, rules-based extraction, model-generated, or mixed workflow;source_object: the file or page actually being summarized;source_role: official, institutional, news, community archive, platform repost, and related roles;generated_or_observed_at: generation time when available, plus an observation time;uncertainty: missing upstream link, unresolved date, parser failure, version ambiguity, and related gaps;reopen_path: the route to full text, PDF, archive capture, or upstream source.
This contract positions an automated summary as a navigation interface. Scholarship on AI and born-digital archives shows that automation can broaden access to large collections while interpretive infrastructure determines how researchers understand machine processing (Jaillant & Caputo 2022). Visible generation status and uncertainty fields allow retrieval speed and archival traceability to reinforce each other.
11. Counterevidence, boundaries, and repeatability
First, search snippets and rank positions are time-dependent. Search systems recrawl pages, refresh indexes, test presentation, and modify snippet generation. Platform pages also change titles, tags, and descriptions. This study therefore treats September 8, 2026 as an observation date. A claim about persistent ranking behavior requires repeated sampling over time. observed_at is a required historical field.
Second, short summaries have positive informational value. They let readers triage large result sets and identify likely relevance quickly. Compression is a descriptive concept here, referring to how a limited interface selects information. High-quality compression can preserve key source signals and a clear path back to the underlying object.
Third, a public search result represents one state of a discovery interface. Public opinion, population attitudes, platform-wide recommendation probabilities, and personalized user experience each require separate evidence. Research on echo chambers and content moderation demonstrates the complex relationship among ranking, interaction, and governance (Gao, Liu & Gao 2023; Jiang et al. 2022). This observatory confines its conclusions to re-openable public pages and observed retrieval representations.
Fourth, historical identity language requires provenance. The 2009 medical rule, 2011 Jin Xing reporting, 2016 Mr C reporting, and current community vocabulary belong to different institutional and temporal settings. Historical titles can remain source-faithful while current editorial text uses stable respectful terminology. identity_provenance allows the two layers to coexist.
Fifth, fluency and evidentiary completeness are separate axes for automated summaries. The PUH3 archive exposes an unresolved original link; the 2017 survey archive exposes a machine-generation notice. Both choices reveal system boundaries to the reader. This produces an auditable form of archival modesty: known fields support discovery, while unresolved fields define the next research step.
12. Practical recommendations for transgender public-history infrastructure
For historians, the lowest-cost practice is to preserve the query frame, observation date, result title or snippet, result URL, and source role. When citing a statistic, add the denominator and analytical subset. When citing litigation, add the procedural stage. When citing identity language, add the speaker or institutional provenance. When citing a video-platform page, record both creator and platform uploader. These fields require little storage and substantially improve later interpretability.
For community archives, machine summaries can continue to support large-scale navigation while summary_status, source_role, original_url_status, and archived_at remain visible to readers. An unresolved upstream link should be represented as an explicit research state. McMurry and colleagues’ work on persistent identifiers emphasizes the relationship among object identity, locatability, and metadata (McMurry et al. 2017). Community archives can apply the same logic at the web-interface level by separating “what is this object?” from “where can it currently be reached?”
For specialist search services, source profiles are a high-value retrieval feature. Results can expose source category, last verification date, region, content type, archival state, and version relation alongside relevance. Users can then distinguish an official page, contemporaneous news report, community copy, and later index before opening the result. TSIndex’s current source archive demonstrates one implementation direction (TSIndex 2026c).
A future platform audit can convert this qualitative provenance study into a longitudinal panel. The same historical queries can run weekly in a fixed environment, saving the first several results’ titles, domains, dates, snippets, source roles, and automated-summary markers. After several months, researchers can measure field retention and source diversity over time. Everyday algorithm auditing and multi-situated platform methods offer methodological foundations for such a design (Shen et al. 2021; Dieter et al. 2019).
Conclusion
Once mainland China’s transgender public history enters search, historical objects commonly reach readers first as compressed representations. A governance rule becomes an eligibility checklist. A survey becomes a number. Several legal tracks become a “victory.” A hospital archive becomes a service summary. A public figure moves between an identity controversy and a hosting role depending on the query frame. A film result can place platform-upload provenance ahead of creator provenance.
The discovery compression chain locates these transformations: historical object → source record → page metadata → index representation → query frame → result snippet/automated summary → reader inference. It treats search results as historical mediators rather than transparent windows. Titles, snippets, tags, and generated descriptions are navigation devices in the present and potential historical artifacts for the future.
The retrieval context ledger supplies a low-cost mechanism for restoring provenance. Recording source role, historical date, institutional actor, denominator, identity-language provenance, legal stage, creator/uploader relation, and summary status lets search remain concise while serious citation reconnects to the evidence structure. For automated summaries, the summary provenance contract adds generation status, source object, uncertainty, and a path back to the underlying record.
The strongest conclusion from this fixed sample is therefore a two-level definition of discoverability. The first level allows a reader to find an object. The second allows a reader to find that object with enough provenance context to interpret it. Public-web search, community archives, and specialist indexes already contain pieces of this second layer: source profiles, archive relations, machine-generation labels, unresolved fields, and verification dates. Turning those pieces into repeatable, durable retrieval metadata offers a practical next step for the digital infrastructure of transgender history.
References
- Poell, Thomas, David Nieborg, and José van Dijck. 2019. “Platformisation.” Internet Policy Review. https://policyreview.info/pdf/policyreview-2019-4-1425.pdf
- DeVito, Michael Ann. 2022. “How Transfeminine TikTok Creators Navigate the Algorithmic Trap of Visibility Via Folk Theorization.” https://dl.acm.org/doi/pdf/10.1145/3555105
- Haimson, Oliver L., Daniel Delmonaco, Peipei Nie, and Andrea Wegner. 2021. “Disproportionate Removals and Differing Content Moderation Experiences.” https://dl.acm.org/doi/pdf/10.1145/3479610
- Shen, Hong, Alicia DeVos, Motahhare Eslami, and Kenneth Holstein. 2021. “Everyday Algorithm Auditing.” https://dl.acm.org/doi/pdf/10.1145/3479577
- Duguay, Stefanie. 2016. “Lesbian, Gay, Bisexual, Trans, and Queer Visibility Through Selfies.” https://journals.sagepub.com/doi/pdf/10.1177/2056305116641975
- MacKinnon, Kinnon R., Hannah Kia, and Ashley Lacombe-Duncan. 2021. “Examining TikTok’s Potential for Community-Engaged Digital Knowledge Mobilization With Equity-Seeking Groups.” https://www.jmir.org/2021/12/e30315/PDF
- Augustaitis, Laima, Leland Merrill, Kristi E. Gamarel, and Oliver L. Haimson. 2021. “Online Transgender Health Information Seeking.” https://dl.acm.org/doi/pdf/10.1145/3411764.3445091
- Jiang, Jialun Aaron, Peipei Nie, Jed R. Brubaker, and Casey Fiesler. 2022. “A Trade-off-centered Framework of Content Moderation.” https://dl.acm.org/doi/pdf/10.1145/3534929
- Simpson, Ellen, Andrew Hamann, and Bryan Semaan. 2022. “How to Tame ‘Your’ Algorithm.” https://dl.acm.org/doi/pdf/10.1145/3492841
- Gao, Yichang, Fengming Liu, and Lei Gao. 2023. “Echo chamber effects on short video platforms.” Scientific Reports. https://www.nature.com/articles/s41598-023-33370-1.pdf
- Bartolome, Ava, and Shuo Niu. 2023. “A Literature Review of Video-Sharing Platform Research in HCI.” https://dl.acm.org/doi/pdf/10.1145/3544548.3581107
- Shang, Zheyu. 2025. “Shifting platform governance: examining participatory content moderation on a Chinese platform Bilibili.” https://www.tandfonline.com/doi/pdf/10.1080/1369118X.2025.2520004?needAccess=true
- Klein, Martin, Herbert Van de Sompel, Robert Sanderson, Harihar Shankar, Lyudmila Balakireva, Ke Zhou, and Richard Tobin. 2014. “Scholarly Context Not Found: One in Five Articles Suffers from Reference Rot.” PLOS ONE. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0115253
- Jones, Shawn M., Herbert Van de Sompel, Harihar Shankar, Martin Klein, Richard Tobin, and Claire Grover. 2016. “Scholarly Context Adrift: Three out of Four URI References Lead to Changed Content.” PLOS ONE. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0167475
- Jaillant, Lise, and Annalina Caputo. 2022. “Unlocking Digital Archives: Cross-Disciplinary Perspectives on AI and Born-Digital Data.” https://link.springer.com/article/10.1007/s00146-021-01367-x
- Brügger, Niels. 2012. “When the Present Web Is Later the Past: Web Historiography, Digital History and Internet Studies.” http://www.ssoar.info/ssoar/handle/document/38378
- Acker, Amelia, and Adam Kreisberg. 2020. “Social Media Data Archives in an API-Driven World.” https://link.springer.com/article/10.1007/s10502-019-09325-9
- Caswell, Michelle, and Marika Cifor. 2016. “From Human Rights to Feminist Ethics: Radical Empathy in the Archives.” http://archivaria.ca/index.php/archivaria/article/view/13557
- Caswell, Michelle, Alda Allina Migoni, Noah Geraci, and Marika Cifor. 2017. “‘To Be Able to Imagine Otherwise’: Community Archives and the Importance of Representation.” https://escholarship.org/uc/item/1h54v9m9
- Wakimoto, Diana K., Christine Bruce, and Helen Partridge. 2013. “Archivist as Activist: Lessons from Three Queer Community Archives in California.” https://eprints.qut.edu.au/58605/
- Fickers, Andreas. 2020. “Update für die Hermeneutik. Geschichtswissenschaft auf dem Weg zur digitalen Forensik?” http://orbilu.uni.lu/handle/10993/44086
- Nix, Adam, and Stephanie Decker. 2021. “Using Digital Sources: The Future of Business History?” https://research-information.bris.ac.uk/en/publications/using-digital-sources-the-future-of-business-history
- Dieter, Michael, Carolin Gerlitz, Anne Helmond, Nathaniel Tkacz, Fernando van der Vlist, and Esther Weltevrede. 2019. “Multi-Situated App Studies: Methods and Propositions.” https://doi.org/10.1177/2056305119846486
- Mandolessi, Silvana. 2024. “The Digital Turn in Memory Studies.” https://doi.org/10.1177/17506980231204201
- McMurry, Julie A., et al. 2017. “Identifiers for the 21st Century.” PLOS Biology. https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.2001414
- 中华人民共和国卫生部. 2009. 《卫生部办公厅关于印发〈变性手术技术管理规范(试行)〉的通知》. https://www.nhc.gov.cn/zwgk/wtwj/201304/5a310ec69d264d4a890949d0b2fbcaf7.shtml
- Trans Chinese Digital Archive. 2026a. “变性手术技术管理规范(试行).” https://digital.transchinese.org/%E6%94%BF%E5%BA%9C%E5%8F%8A%E5%AE%98%E6%96%B9%E7%BB%84%E7%BB%87%E6%96%87%E4%BB%B6/%E4%B8%AD%E5%9B%BD%E5%A4%A7%E9%99%86/%E5%8F%98%E6%80%A7%E6%89%8B%E6%9C%AF%E6%8A%80%E6%9C%AF%E7%AE%A1%E7%90%86%E8%A7%84%E8%8C%83%EF%BC%88%E8%AF%95%E8%A1%8C%EF%BC%89/
- Project Trans. 2026. “变性手术技术管理规范(试行).” https://project-trans.org/china-legal/spec/2009-11-13/srs/readme/
- China Development Brief / Beijing LGBT Center. 2017. “中国跨性别群体生存现状调查报告圆满发布.” https://www.chinadevelopmentbrief.org.cn/news/detail/17772.html
- Beijing LGBT Center & Department of Sociology, Peking University. 2017. 2017中国跨性别群体生存现状调查报告. https://cnlgbtdata.com/doc/43/
- TGR Transgender Resource Center. 2017. “2017中国跨性别群体生存现状调研报告.” https://www.tgr.org.hk/index.php/zh/china/china-info/378-2017%E4%B8%AD%E5%9B%BD%E8%B7%A8%E6%80%A7%E5%88%AB%E7%BE%A4%E4%BD%93%E7%94%9F%E5%AD%98%E7%8E%B0%E7%8A%B6%E8%B0%83%E7%A0%94%E6%8A%A5%E5%91%8A
- Trans Chinese Digital Archive. 2026b. “2017中国跨性别群体生存现状调查报告_北京同志中心.” https://digital.transchinese.org/%E6%9D%82%E5%BF%97%E5%8F%8A%E6%96%B0%E9%97%BB%E6%8A%A5%E9%81%93/%E4%B8%AD%E5%9B%BD%E5%A4%A7%E9%99%86/2017%E4%B8%AD%E5%9B%BD%E8%B7%A8%E6%80%A7%E5%88%AB%E7%BE%A4%E4%BD%93%E7%94%9F%E5%AD%98%E7%8E%B0%E7%8A%B6%E8%B0%83%E6%9F%A5%E6%8A%A5%E5%91%8A%E5%8F%91%E5%B8%83_%E5%8C%97%E4%BA%AC%E5%90%8C%E5%BF%97%E4%B8%AD%E5%BF%83_page/
- 中国新闻网. 2016. 《“女扮男装”员工被开除 当事人申请仲裁已获立案》. https://www.chinanews.com.cn/m/sh/2016/03-16/7799284.shtml
- Radio Free Asia. 2017a. “首名跨性人告僱主勝訴獲償 續提控要求道歉.” https://www.rfa.org/cantonese/news/discrimination-01022017081923.html
- Radio Free Asia. 2017b. “贵阳法院就跨性别者就业歧视案做出判决 原告‘C’先生获胜.” https://www.rfa.org/mandarin/Xinwen/6-07272017134541.html
- Common Language. 2020. “影响性诉讼⑦:这次,把‘性别认同及性别表达’写进判决.” https://commonlanguage.github.io/TYarchives2020/20200518_1%E5%BD%B1%E5%93%8D%E6%80%A7%E8%AF%89%E8%AE%BC%E2%91%A6%E8%BF%99%E6%AC%A1%EF%BC%8C%E6%8A%8A%E2%80%9C%E6%80%A7%E5%88%AB%E8%AE%A4%E5%90%8C%E5%8F%8A%E6%80%A7%E5%88%AB%E8%A1%A8%E8%BE%BE%E2%80%9D%E5%86%99%E8%BF%9B%E5%88%A4%E5%86%B3.html
- PUH3 Department of Plastic Surgery. 2025. “北医三院易性症综合诊疗就诊指南.” https://www.puh3.net.cn/zxwk/info/1005/4211.htm
- PUH3 Department of Endocrinology. 2026. “北医三院性别医学团队在JAMA Network Open发表跨性别相关研究成果.” https://www.puh3.net.cn/nfmk/info/1271/4602.htm
- PUH3 Department of Plastic Surgery. 2018. “北医三院易性症序列医疗团队成立.” https://www.sar.com.cn/xinwen/news/12295.html
- Trans Chinese Digital Archive. 2026c. “北医三院‘易性症’序列治疗就诊指南.” https://digital.transchinese.org/%E7%A4%BE%E7%BE%A4%E5%8F%8ANGO%E6%96%87%E4%BB%B6/%E5%8C%BB%E9%99%A2%E5%92%8C%E5%8C%BB%E7%96%97%E4%BD%93%E7%B3%BB/%E5%8C%97%E5%8C%BB%E4%B8%89%E9%99%A2%E2%80%9C%E6%98%93%E6%80%A7%E7%97%87%E2%80%9D%E5%BA%8F%E5%88%97%E6%B2%BB%E7%96%97%E5%B0%B1%E8%AF%8A%E6%8C%87%E5%8D%97_page/
- China News. 2011a. “金星自曝因变性被封杀 愤慨斥遭歧视:决不接受.” https://www.chinanews.com/yl/2011/09-22/3346646.shtml
- China News. 2011b. “金星不满因‘变性’被取消评委资格 宋丹丹力挺.” https://www.chinanews.com/yl/2011/09-23/3349842.shtml
- China News. 2011c. “舞蹈家金星谈‘封杀’:这实在是太滑稽了.” https://www.chinanews.com.cn/cul/2011/09-30/3365004.shtml
- Sina Entertainment. 2012. “脱口秀《金星撞火星》17日首播 金星担任主持人.” https://ent.sina.com.cn/v/m/2012-03-07/13393574628.shtml
- China News. 2015. “金星推脱口秀叫板崔永元.” https://www.chinanews.com.cn/yl/2015/01-27/7007840.shtml
- People’s Daily. 2015. “东方卫视推出《金星秀》 舞蹈家想靠嘴征服世界.” https://media.people.com.cn/n/2015/0206/c40606-26518070.html
- Yennefer Fang. 2021. “A Day of Trans.” Vimeo. https://vimeo.com/647384537
- Bilibili. 2022. “【4K】跨越性别的一天 A Day of Trans.” https://www.bilibili.com/video/BV1pD4y1e7KB/
- TSIndex. 2026a. “多元性别搜索引擎.” https://www.tsindex.org/
- TSIndex. 2026b. “About / 关于我们.” https://www.tsindex.org/about
- TSIndex. 2026c. “来源档案.” https://www.tsindex.org/archive/
- Memento. 2026. “Time Travel for the Web.” https://mementoweb.org/about/