Abstract
A large share of mainland China’s recent transgender history survives as born-digital documents. National medical rules appear as government pages and attachments. Community surveys circulate as PDFs. Hospitals record the formation of clinical teams in news pages. Legal advocates reorganize litigation and regulation into topic pages. Digital archives then add author, year, original URL, file size, format, and checksum metadata to documents whose original distribution paths have changed. In 2026, when AI-assisted retrieval and text analysis increasingly sit inside ordinary research workflows, the public visibility of a source can be separated into more specific technical properties: whether an interface identifies the object clearly, whether its text can be parsed, whether headings and pages provide structural anchors, whether provenance relations are explicit, whether the file has a stable identity, and whether multiple resolvable entry points lead to the same historical object.
This article observes twelve groups of public sources: the 2009 national technical rule for “sex reassignment surgery”; the 2016 national LGBTI survey; the 2017 national transgender survey and the school/document-change report; the 2018 legal-gender-recognition review and Peking University Third Hospital team page; the 2019 document-change manual and Amnesty International health-care report; the 2021 National Transgender Health Survey; the 2022 G05 gender-reassignment-technology rule; and legal and community archival mirrors. The unit of observation is the public interface through which a researcher, search engine, or text-processing system reaches each source in September 2026. The article therefore studies document infrastructure and evidence interfaces.
I propose a machine-readable evidence ladder: object identity → parseable text → structural anchors → provenance relations → file identity → redundant discovery entrances. The six layers answer, respectively, what the object is, where its text lives, how sections can be located, where it came from, which file/version is being processed, and which independent interfaces lead back to it. I also propose AI discovery debt, a way to describe the transformation work still required before a public historical object becomes citation-ready structured evidence. A file-only entrance carries metadata-recovery work; an image-based document carries text-recovery work; an automatically generated archive summary requires return to the source document for evidentiary checking; an object with a descriptive HTML landing page, extractable text, a stable file, and explicit provenance relations carries a shorter transformation path and therefore lower discovery debt.
The fixed sample shows a set of complementary infrastructures, each optimized for different historical tasks. Government HTML preserves authority and publication context. Institutional landing pages connect descriptive metadata to multilingual PDFs. Community archives add file-level metadata and original-source relations. Legal topic sites turn long regulations into addressable web structure. PDFs retain page numbers, visual layout, figures, and tables; HTML and archival metadata provide discovery, section targeting, and source connections. An AI-ready historical infrastructure can therefore preserve original files while also maintaining machine-readable identity, text, structure, provenance, and version relations.
Keywords: transgender history in mainland China; machine readability; PDF; digital archives; artificial intelligence; information retrieval; provenance; source criticism
1. Research question: from “the page survives” to “how a machine reaches the evidence”
Digital-history research has made link decay, page versions, web archives, and retrieval paths part of historical method. Klein and colleagues showed how scholarly references suffer reference rot; Jones and colleagues demonstrated that a URI can remain live while the content to which it points changes; the Memento framework connects original URLs to time-specific archived representations (Klein et al. 2014; Jones et al. 2016; Sanderson et al. 2011). GenderLibs’ September 1 observatory used this tradition to describe a source-survival stack that connects original publication, contemporaneous replicas, community archives, general web archives, discovery indexes, and scholarly citation.
The present study moves one level inside the document. A historical object may have a stable URL while appearing in several forms: full HTML, an HTML landing page linked to a PDF, a standalone PDF, an archival metadata page linked to a preserved file, or a normalized legal page that reproduces a regulation in sectioned web text. Human readers can use each of these forms. Search engines, corpus tools, and generative systems must turn them into titles, bodies, sections, tables, links, file identities, and provenance relations. Lara Putnam’s account of the “text-searchable” archive shows how full-text retrieval changes the questions historians can ask and the geographical connections they can make. Research on archives and AI similarly treats collection organization, description, and interfaces as part of the conditions under which automated discovery operates (Putnam 2016; Colavizza et al. 2022).
The central question is therefore: which machine-readable interfaces do public sources on mainland China’s transgender history expose in 2026, what historical value does each document form preserve, and how can researchers record those interfaces so that AI-assisted extraction remains traceable to evidence?
The article’s answer is that machine readability is layered. A text layer opens sentences to retrieval. A structure layer exposes sections, pages, tables, and captions. A provenance layer explains the relation between an original publication, a mirror, an archive record, and a later description. A file-identity layer distinguishes versions. A discovery layer creates paths from query language to the object. Durable historical infrastructure connects these layers so that human reading, full-text search, automated extraction, and citation share the same source coordinates.
2. Materials and method: a fixed sample and six observable layers
The observation date is September 15, 2026. Sources enter the sample when they meet three conditions: they concern mainland China’s transgender history directly; they have a public, resolvable web entrance; and together they represent distinct publishing traditions across policy, surveys, medicine, legal advocacy, and community knowledge. The twelve source groups come from national health authorities, UNDP, China Development Brief, Peking University Third Hospital, Amnesty International, CNLGBTDATA, the Chinese Transgender Digital Archive, Project Trans, and Common Language.
The method records technical facts visible through public interfaces. UNDP’s 2016 survey page, for example, supplies a title, date, implementing organizations, description, and download path, while the associated publication interface exposes Chinese and English PDFs. Amnesty International’s 2019 report has a report landing page, language selection, and a PDF whose table of contents and body text can be extracted. The archive record for the 2017 transgender survey exposes a filename, format, byte size, MD5 checksum, archive date, original link, author, region, year, tags, and a downloadable file (UNDP 2016; Amnesty International 2019a; Chinese Transgender Digital Archive 2026a).
I divide those observations into six layers:
| Layer | Observable fields | Historical use |
|---|---|---|
| Object identity | title, publisher/author, date, document number | establish the cited object and institutional actor |
| Parseable text | HTML body, PDF text layer, extractable paragraphs | full-text retrieval, quotation targeting, corpus work |
| Structural anchors | headings, contents, tables, page numbers, labels | section-level verification and comparison |
| Provenance relations | original link, mirror relation, archive note | separate publication, copy, archive, and redescription |
| File identity | filename, format, size, checksum, version | distinguish copies and versions |
| Redundant discovery entrances | institutional page, archive record, legal mirror, catalog | provide multiple paths to the same object |
This approach connects to several strands of document research. OCR and post-OCR scholarship studies the recovery and correction of characters. Historical named-entity-recognition research studies people, organizations, locations, and dates extracted from historical text. Archives-and-AI research studies discovery, description, and access workflows shaped by machine learning (Nguyen et al. 2021; Ehrmann et al. 2023; Muehlberger et al. 2019; Colavizza et al. 2022). The present article places those questions inside a specific Chinese-language transgender-history corpus and asks how publication interfaces affect later evidence work.
3. HTML as an object-identity interface: the 2009 national rule and a 2018 hospital record
The 2009 Technical Management Specification for Sex Reassignment Surgery (Trial) retains an institutional publication context on the national health-authority website, while Project Trans presents the same rule as a normalized topic page with section headings and addressable full text (Ministry of Health 2009; Project Trans 2026a). The two pages perform different historical functions. The government page anchors authority and publication context. The legal topic page improves section-level navigation and places the text inside a related regulatory chronology. A researcher can record both and retain both institutional provenance and retrieval convenience.
HTML is especially useful when object identity and body text share the same page structure. A title can enter a search index. Paragraphs can enter full-text retrieval. Headings can become semantic anchors. Hyperlinks can relate the current text to a parent notice, a later version, or a neighboring rule. Putnam’s analysis of digitized, text-searchable sources emphasizes the shift from catalog-level discovery into cross-document searching inside the sources themselves; that shift is highly relevant to policy terminology, where terms such as “sex reassignment surgery,” “gender reassignment technology,” “transsexualism,” and “gender incongruence” can be located by regulatory version (Putnam 2016).
The Peking University Third Hospital page announcing the creation of a multidisciplinary team in 2018 provides another HTML object. It preserves a publication date, hospital-department source, title, and continuous prose describing psychology, endocrinology, andrology, plastic surgery, and related services (PUH3 Department of Plastic Surgery 2018). Its historical value comes from institutional self-description at a dated moment. A text system can extract institution, date, service categories, and terminology, then compare them with national regulation or clinical literature.
The source role remains essential. A government notice, a hospital news item, and a normalized legal mirror may all use HTML, yet they represent official publication, institutional memory, and thematic republication. The provenance-relation layer preserves that difference after all three have been converted into searchable text.
4. Landing page plus PDF: UNDP’s 2016 and 2018 dual interfaces
UNDP’s 2016 Being LGBTI in China illustrates a mature two-layer publication interface. Its public landing page exposes the report title, May 16, 2016 date, survey scope, implementing organizations, and a download path. A related UNDP publication interface explicitly offers English and Chinese PDF files (UNDP 2016). The landing page is well suited to object discovery and descriptive indexing; the PDF preserves the complete report, pagination, visual organization, tables, and figures.
The 2018 Legal Gender Recognition in China: A Legal and Policy Review uses a comparable structure. Its UNDP page lists Chinese and English PDFs before the descriptive body and preserves date, purpose, and regional-program context. A contemporaneous press release adds launch-event context, partners, and participant statements (UNDP 2018a, 2018b). The researcher therefore encounters three evidentiary objects: descriptive metadata on the report page, the long-form publication itself, and event-level context in the release announcement.
This division of labor is useful for AI-assisted work. Object resolution begins with HTML metadata; long-form analysis moves into the PDF text layer; pagination remains available for conventional citation. Westergaard and colleagues’ comparison of text mining across full texts and abstracts shows the richer relation and terminology evidence available in full documents, while Starr and colleagues’ discussion of human- and machine-accessible cited data emphasizes identifiers, location, and machine reachability as conditions of reuse (Westergaard et al. 2018; Starr et al. 2015).
PDF also preserves a visual historical record. Tables, page breaks, captions, appendices, and bilingual layout may themselves matter. A strong digital-history interface therefore keeps the PDF intact while adding or retaining an indexable HTML identity page. The arrangement supports visual evidence and machine discovery at the same time.
5. The 2017 national transgender survey: publication page, downloadable file, and archival metadata
The 2017 national transgender survey now has several public entrances. China Development Brief’s release article preserves a title, source, author, date, launch context, and a download link; a separate publication record registers the report as a standalone publication. The Transgender Resource Center Hong Kong also retains a report page, creating a contemporaneous cross-regional discovery route (China Development Brief 2017a, 2017b; Transgender Resource Center Hong Kong 2017).
The Chinese Transgender Digital Archive adds file-level description. Its current record identifies Chinese_Transgender_Population_General_Survey_Report.pdf, labels the PDF format, records file size, MD5 checksum, archive date, original URL, author, region, year, and tags, and provides a downloadable copy (Chinese Transgender Digital Archive 2026a). These fields turn a report into a computational archival object. Filename and checksum aid duplicate recognition; the original URL preserves the source relation; author and year support catalog retrieval; tags add topic-level entrances.
The record also labels its descriptive summary and supplemental information as automatically generated for retrieval and reference. That label is valuable provenance in its own right. It tells a researcher that the summary belongs to the archive’s discovery layer while the report file remains the substantive evidence layer. An AI-oriented archive can express the distinction in fields such as generated_description=true, evidence_source=file, and original_url=..., enabling downstream systems to choose evidence layers deliberately.
Gil and colleagues’ recommendations for documenting research from data and software through provenance offer a broader parallel: reusable objects become easier to audit when their transformations and relations are explicit (Gil et al. 2016). In transgender-history research, file-level provenance also helps future investigators decide whether two copies separated by a site migration represent the same version.
6. A 2017 school/document report and a 2019 document-change manual: community documents gain machine entrances
The Transgender School and Document-Change Survey Report and the Transgender Document-Change Manual represent two kinds of community and legal-practice document. The former joins freedom-of-information work, a survey of school experiences, and document-change issues. The latter organizes procedures for changing identity, household-registration, education, and professional documents into a service manual. Their original historical value comes from advocacy and service practice; later archival interfaces add a second value as research infrastructure.
The Chinese Transgender Digital Archive currently provides a dedicated metadata page and PDF download for the school/document report. Its record for the document-change manual supplies the title, the author label “Rainbow Lawyers,” a 2019 date, PDF filename, byte size, MD5 checksum, archive date, and original link (Chinese Transgender Digital Archive 2026b, 2026c). This archival wrapper makes isolated files searchable as objects. A machine can inspect the HTML metadata layer to identify topic, date, and creator before opening the document body.
The wrapper also preserves migration history. When an organization website, file host, or document server changes, an archive record can retain original_link and a later archive date. Acker and Kreisberg’s study of API-driven social-media archives treats interfaces as archival conditions, while Caswell and Cifor place relationship and responsibility inside archival description (Acker & Kreisberg 2020; Caswell & Cifor 2016). Community history benefits from both ideas: technical fields show where a file came from; descriptive practice explains who created it and what kind of evidence it contains.
This design also helps with terminology change. An older file title can remain intact while contemporary subject tags create new discovery paths. The result is a bridge that keeps historical language and current search vocabulary simultaneously addressable.
7. Amnesty International’s 2019 report: a text-bearing PDF as a deep-reading interface
Amnesty International’s 2019 report on barriers to gender-affirming care in China has a layered publication design. The report landing page gives title, date, document index, language selection, and a PDF download. A Simplified Chinese landing page exposes the corresponding edition. The English PDF preserves a cover, publication information, table of contents, numbered sections, glossary, and full body text (Amnesty International 2019a, 2019b, 2019c).
Public parsing of the forty-seven-page PDF recovers the contents as section names and page numbers and extracts body paragraphs page by page. For an AI workflow, that means the document has already crossed the most basic image-to-character threshold. The next task becomes structural recovery: distinguishing body text from line wraps, headers, notes, tables, and page-level artifacts while maintaining page coordinates. Baviskar and colleagues’ review of AI processing for unstructured documents, together with Nguyen and colleagues’ post-OCR survey, shows why character availability and document understanding belong to separate stages of a pipeline (Baviskar et al. 2021; Nguyen et al. 2021).
The PDF contributes stable page coordinates, visual hierarchy, and footnote placement as a distinct evidence layer. The landing page contributes a stronger discovery and identity interface. The machine-readable evidence ladder therefore records parseable text separately from structural anchors. A PDF may have excellent text extraction and still require layout normalization; an HTML page may have clear headings and links while using flow-based location instead of fixed pagination. Their strengths combine.
8. The 2021 National Transgender Health Survey: large-file identity and two catalog interfaces
The 2021 National Transgender Health Survey adds file scale to the problem. CNLGBTDATA’s catalog record identifies the report, publication date, major creators, and subject classification; the site also preserves a long-form introduction and a public PDF. The Chinese Transgender Digital Archive records a filename, PDF format, a file size of roughly fifty-four megabytes, MD5 checksum, archive date, original link, region, year, and tags (CNLGBTDATA 2023a, 2023b; Chinese Transgender Digital Archive 2026d).
These entrances show how even “large file” can become useful machine metadata. Before downloading the full document, a system can know what the object is, when it was produced, how large the current file is, and which archived or original entrance it belongs to. At corpus scale, the HTML catalog layer can build an object inventory first; full files can then be fetched selectively for tasks that need them. Digital-humanities infrastructure has long stressed the relation between collection metadata and full-text access. Borgman’s call for humanities cyberinfrastructure and Terras’s work on open digitized collections both connect forms of access to the kinds of research that become feasible (Borgman 2009; Terras 2015).
The archive record’s checksum serves another role. A checksum answers a version question precisely: whether two mirrored downloads are byte-identical, whether a migration replaced the file, or whether a local cache matches the catalog record. An AI index can use file identity to avoid embedding the same version repeatedly and to trigger reprocessing when a file changes.
9. The 2022 G05 rule: policy readability across notice and attachment
In 2022, the National Health Commission issued the national restricted-technology directory and clinical-application rules. The notice page preserves the document number, publication date, effective relationship to earlier rules, and attachment entrances. Gender reassignment technology appears as G05 in the updated restricted-technology framework (National Health Commission 2022). Historical analysis therefore requires both the notice object and the attached standard object.
The notice supplies version relations. It says who issued the material, when the change took effect, which earlier regulatory set was superseded, and how the attachment belongs to the formal notice. The attachment supplies the specific technical provisions. A system that captures only an attachment body can lose effective-date and supersession context; a system that captures only a notice summary can lose the clauses. A complete policy object is better represented relationally: notice -> has_attachment -> G05 standard, with publication date and supersedes links attached to the appropriate nodes.
Fickers’s account of historical scholarship moving toward digital forensics highlights the value of inspecting digital objects, versions, and conditions of production (Fickers 2020). Regulatory attachments illustrate why provenance relations belong inside machine readability. Sustainable policy retrieval requires both clause text and clause identity; a provenance graph therefore operates alongside a full-text index as a coequal layer.
10. The machine-readable evidence ladder as an auditable object model
The twelve source groups can be summarized in a reusable record:
{object_id, title, publisher, published_at, source_role, landing_url, file_url, text_mode, structure_mode, original_url, archive_url, filename, format, size, checksum, language, version, relation_to_other_object, observed_at}
text_mode can record states such as html, pdf_text, ocr_text, or image. structure_mode can describe headings, contents, page numbers, tables, and captions. source_role can distinguish official publication, institutional record, community report, archive description, legal mirror, and press release. Each value should come from a public object that can be rechecked.
The machine-readable evidence ladder turns the record into six consecutive research tasks:
- Object identity: title, creator or issuing institution, date, and document number establish what is being processed.
- Parseable text: HTML or a document text layer opens terms, sentences, and paragraphs to retrieval.
- Structural anchors: sections, contents, page numbers, tables, and heading hierarchies make extraction verifiable.
- Provenance relations: original links, mirror notes, archive notes, and parent notices preserve historical context.
- File identity: filename, size, checksum, and version distinguish similarly titled objects and later replacements.
- Redundant discovery entrances: institutional pages, archive records, legal mirrors, catalogs, and search surfaces provide multiple ways back to the object.
This ladder works at a different scale from the source-survival stack introduced in the September 1 observatory. The source-survival stack asks how an object persists across sites and time. The machine-readable evidence ladder asks which auditable fields become available once the object reaches a research workflow. The two models connect naturally: preservation layers keep the object alive; machine-readable layers keep its identity, structure, and provenance usable during extraction.
11. AI discovery debt: measuring transformation work before evidence becomes citation-ready
I define AI discovery debt as the transformation work between a public source entrance and citation-ready structured evidence. The debt can come from object metadata that requires recovery, text confined to images, section structure that requires reconstruction, origin relations that require source criticism, version identity that requires reconstruction, or a single fragile entrance. The concept records work and evidentiary friction across source interfaces.
A landing page that provides title, date, publisher, summary, and bilingual PDFs lets a system create an object record immediately and fetch full text as a second step. An archival record with filename, checksum, original URL, and tags provides provenance before text processing begins. A standalone PDF still preserves a complete readable object, while a system must derive title, author, date, and section structure from inside the file. An image-based document adds OCR and layout reconstruction. Research on handwritten-text recognition and post-OCR correction shows that each transformation stage produces identifiable recognition and correction tasks (Muehlberger et al. 2019; Nguyen et al. 2021).
Discovery debt can be decomposed operationally:
discovery_debt = metadata_recovery + text_recovery + structure_recovery + provenance_recovery + version_recovery
This is a workflow model built from component states. A research team can label each component ready, derived, or manual_review. When an HTML page supplies date and title directly, metadata_recovery=ready. When body text comes from OCR, text_recovery=derived. When the original publishing institution requires source criticism, provenance_recovery=manual_review. The state model captures historical uncertainty more faithfully than a single readable/unreadable switch.
12. Counterevidence and limits: machine readability and historical authenticity perform different jobs
Machine readability increases retrieval and extraction capacity. Historical authenticity continues to depend on source criticism. The Chinese Transgender Digital Archive’s automatically generated descriptions help discovery, while the record itself marks them as generated reference material and keeps the underlying file available as evidence (Chinese Transgender Digital Archive 2026a). A normalized legal mirror can provide excellent section navigation while an official publication page anchors institutional authority. Used together, the two interfaces provide a stronger evidentiary chain.
PDF demonstrates the same complementarity. Stable page numbers and visual layout support precise citation and the history of presentation; HTML supports full-text discovery and link relations. Scanned images preserve marks on a page; OCR adds a search layer. Ehrmann and colleagues’ survey of named-entity recognition in historical documents shows how language variation, layout, and recognition errors affect automated entity extraction. Colavizza and colleagues place AI opportunities in archives alongside description, bias, transparency, and professional judgment (Ehrmann et al. 2023; Colavizza et al. 2022).
Sample scale is another boundary. This is a fixed-sample observatory covering representative policy, survey, hospital, legal, and community document forms. Its purpose is to define measurable interfaces and record-keeping fields. A later corpus-scale study can measure what proportion of a collection has text layers, how OCR error varies by period and layout, or which fields most improve bilingual retrieval. The fixed sample supplies the measurement unit needed before those larger statistics become meaningful.
13. Practical design for the next generation of archives and AI retrieval
The first practice is to give every important file an HTML object-identity page. That page should retain original title, creator or issuing institution, original publication date, language, file link, source role, and original URL. The file remains intact; HTML carries discovery and object relations.
The second practice is to treat extracted text as a derived object with provenance. OCR, PDF text extraction, and human transcription can each produce a retrieval layer while recording tool, date, and source-file checksum. A search hit can then be traced to a specific file and page. The post-OCR literature shows why a text layer itself has a processing history worth retaining (Nguyen et al. 2021).
The third practice is structural preservation. Section headings, page numbers, table identifiers, captions, and footnote relations are more useful to deep retrieval than an undifferentiated stream of text. Historical NER, document understanding, and generative question answering all benefit from those anchors. A response system can return object_id + section + page + source_url together with extracted text.
The fourth practice is provenance as a graph. A UNDP landing page and its PDF, a national notice and its attachment, an archive record and its original URL, or a legal mirror and the formal rule can be connected by explicit relations. An AI system then knows whether a piece of text is an original publication, contemporaneous report, later archival description, or normalized mirror before it summarizes the contents.
The fifth practice is periodic object and file verification. Landing-page state, original links, mirrors, file checksums, and archive entrances can become a light maintenance loop. Massicotte and Botter’s repository case study shows that reference rot remains a maintenance concern even in professional repositories (Massicotte & Botter 2017). The same check can sit inside a digital-preservation workflow.
Conclusion: AI-era historical infrastructure preserves readable relations as well as files
Public sources on mainland China’s transgender history from 2009 through 2022 appear on the 2026 web in many document forms. National health authorities and hospitals provide institutional HTML. UNDP and Amnesty International connect descriptive landing pages to long-form PDFs. China Development Brief preserves publication context and download entrances. CNLGBTDATA provides catalog objects. The Chinese Transgender Digital Archive adds format, size, checksum, original link, and archive date to migrated files. Project Trans turns policy texts into normalized, sectioned web documents. Together, these interfaces form a practical machine-readable historical infrastructure.
The machine-readable evidence ladder separates that infrastructure into six layers: object identity → parseable text → structural anchors → provenance relations → file identity → redundant discovery entrances. The layers preserve object resolution, text, structure, provenance, version identity, and retrieval routes. AI discovery debt records the metadata, text, structure, provenance, and version work that remains before a document becomes citation-ready evidence.
The most useful direction for historical research is therefore to preserve original files and machine-readable relations together. PDF retains page coordinates and visual form. HTML provides object identity and discovery. OCR or text extraction provides full-text access. Archival metadata supplies provenance and file identity. Independent mirrors provide redundant entrances. A future retrieval system can then answer both a content question and a source question: which file contains the statement, which version, which page or section, which publishing institution, which original URL, and which public interfaces still lead to it. That infrastructure lets AI accelerate discovery while keeping historical evidence auditable, reproducible, and traceable.
References
- Putnam, Lara. 2016. “The Transnational and the Text-Searchable: Digitized Sources and the Shadows They Cast.” American Historical Review 121(2): 377–402. https://academic.oup.com/ahr/article-pdf/121/2/377/25736337/zah377.pdf
- Klein, Martin, Herbert Van de Sompel, Robert Sanderson, Harihar Shankar, Lyudmila Balakireva, Ke Zhou, and Richard Tobin. 2014. “Scholarly Context Not Found: One in Five Articles Suffers from Reference Rot.” PLOS ONE. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0115253
- Jones, Shawn M., Herbert Van de Sompel, Harihar Shankar, Martin Klein, Richard Tobin, and Claire Grover. 2016. “Scholarly Context Adrift: Three out of Four URI References Lead to Changed Content.” PLOS ONE. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0167475
- Sanderson, Robert, Mark Edward Phillips, and Herbert Van de Sompel. 2011. “Analyzing the Persistence of Referenced Web Resources with Memento.” https://arxiv.org/abs/1105.3459
- Dougherty, Meghan, Eric T. Meyer, Christine Madsen, Charles van den Heuvel, Arthur Thomas, and Sally Wyatt. 2017. “Researcher Engagement with Web Archives: State of the Art.” https://ecommons.luc.edu/cgi/viewcontent.cgi?article=1014&context=communication_facpubs
- Acker, Amelia, and Adam Kreisberg. 2020. “Social Media Data Archives in an API-Driven World.” https://link.springer.com/article/10.1007/s10502-019-09325-9
- Jaillant, Lise, and Annalina Caputo. 2022. “Unlocking Digital Archives: Cross-Disciplinary Perspectives on AI and Born-Digital Data.” https://link.springer.com/article/10.1007/s00146-021-01367-x
- Brügger, Niels. 2012. “When the Present Web Is Later the Past: Web Historiography, Digital History and Internet Studies.” http://www.ssoar.info/ssoar/handle/document/38378
- Colavizza, Giovanni, Tobias Blanke, Charles Jeurgens, and Julia Noordegraaf. 2022. “Archives and AI: An Overview of Current Debates and Future Perspectives.” https://dl.acm.org/doi/10.1145/3479010
- Ehrmann, Maud, Ahmed Hamdi, Elvys Linhares Pontes, Matteo Romanello, and Antoine Doucet. 2023. “Named Entity Recognition and Classification in Historical Documents: A Survey.” https://dl.acm.org/doi/10.1145/3604931
- Nguyen, Thi Tuyet Haï, Adam Jatowt, Mickaël Coustaty, and Antoine Doucet. 2021. “Survey of Post-OCR Processing Approaches.” https://dl.acm.org/doi/10.1145/3453476
- Muehlberger, Guenter, Louise Seaward, Melissa Terras, et al. 2019. “Transforming Scholarship in the Archives through Handwritten Text Recognition.” https://infoscience.epfl.ch/record/270774/files/10-1108_JD-07-2018-0114.pdf
- Terras, Melissa. 2015. “Opening Access to Collections: The Making and Using of Open Digitised Cultural Content.” https://www.research.ed.ac.uk/files/46486196/Terras2015OIROpeningAccess.pdf
- Borgman, Christine L. 2009. “The Digital Future Is Now: A Call to Action for the Humanities.” https://dhq-static.digitalhumanities.org/pdf/000077.pdf
- Westergaard, David, Hans-Henrik Stærfeldt, Christian Tønsberg, Lars Juhl Jensen, and Søren Brunak. 2018. “A Comprehensive and Quantitative Comparison of Text-Mining in 15 Million Full-Text Articles versus Their Corresponding Abstracts.” https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005962
- Baviskar, Dipali, Swati Ahirrao, Vidyasagar Potdar, and Ketan Kotecha. 2021. “Efficient Automated Processing of the Unstructured Documents Using Artificial Intelligence: A Systematic Literature Review and Future Directions.” https://ieeexplore.ieee.org/document/9402739
- Starr, Joan, Eleni Castro, Mercè Crosas, et al. 2015. “Achieving Human and Machine Accessibility of Cited Data in Scholarly Publications.” https://ir.cwi.nl/pub/25200/pj4237.pdf
- Gil, Yolanda, Cédric H. David, İbrahim Demir, et al. 2016. “Toward the Geoscience Paper of the Future: Best Practices for Documenting and Sharing Research from Data to Software to Provenance.” https://onlinelibrary.wiley.com/doi/10.1002/2015EA000136
- Massicotte, Mia, and Kathleen Botter. 2017. “Reference Rot in the Repository: A Case Study of Electronic Theses and Dissertations (ETDs) in an Academic Library.” https://ejournals.bc.edu/index.php/ital/article/view/9598
- Fickers, Andreas. 2020. “Update für die Hermeneutik. Geschichtswissenschaft auf dem Weg zur digitalen Forensik?” http://orbilu.uni.lu/handle/10993/44086
- Caswell, Michelle, and Marika Cifor. 2016. “From Human Rights to Feminist Ethics: Radical Empathy in the Archives.” http://archivaria.ca/index.php/archivaria/article/view/13557
- Caswell, Michelle, Alda Allina Migoni, Noah Geraci, and Marika Cifor. 2017. “‘To Be Able to Imagine Otherwise’: Community Archives and the Importance of Representation.” https://escholarship.org/uc/item/1h54v9m9
- Guo, Shaohua. 2020. The Evolution of the Chinese Internet: Creative Visibility in the Digital Public. Stanford University Press. https://www.sup.org/books/asian-studies/evolution-chinese-internet
- Ministry of Health of the People’s Republic of China. 2009. “Technical Management Specification for Sex Reassignment Surgery (Trial).” https://www.nhc.gov.cn/zwgk/wtwj/201304/5a310ec69d264d4a890949d0b2fbcaf7.shtml
- National Health Commission of the PRC. 2022. “国家卫生健康委办公厅关于印发国家限制类技术目录和临床应用管理规范(2022年版)的通知.” https://www.nhc.gov.cn/yzygj/c100068/202204/2655831f6f444b00b3e50604e67531f5.shtml
- Project Trans. 2026a. “Technical Management Specification for Sex Reassignment Surgery (Trial).” https://project-trans.org/china-legal/spec/2009-11-13/srs/readme/
- UNDP. 2016. Being LGBTI in China. https://www.undp.org/china/publications/being-lgbti-china
- UNDP. 2018a. Legal Gender Recognition in China: A Legal and Policy Review. https://www.undp.org/china/publications/legal-gender-recognition-china-legal-and-policy-review
- UNDP China. 2018b. “UNDP and China Women’s University Release Legal Gender Recognition Report.” https://www.undp.org/china/press-releases/undp-and-china-womens-university-release-legal-gender-recognition-report
- China Development Brief. 2017a. “中国跨性别群体生存现状调查报告圆满发布.” https://www.chinadevelopmentbrief.org.cn/news/detail/17772.html
- China Development Brief. 2017b. “2017 Chinese Transgender Population General Survey Report.” https://chinadevelopmentbrief.org/publications/2017-chinese-transgender-population-general-survey-report/
- Chinese Transgender Digital Archive. 2026a. “Chinese_Transgender_Population_General_Survey_Report.” https://digital.transchinese.org/%E7%A4%BE%E7%BE%A4%E5%8F%8ANGO%E6%96%87%E4%BB%B6/%E7%BB%9F%E8%AE%A1%E6%8A%A5%E5%91%8A/Chinese_Transgender_Population_General_Survey_Report_page/
- Transgender Resource Center Hong Kong. 2017. “2017中国跨性别群体生存现状调查报告.” https://www.tgr.org.hk/index.php/zh/2017/11/20/2017-chinese-transgender-population-general-survey-report/
- Chinese Transgender Digital Archive. 2026b. “跨性别校园与证件修改调查报告.” https://digital.transchinese.org/%E7%A4%BE%E7%BE%A4%E5%8F%8ANGO%E6%96%87%E4%BB%B6/%E7%BB%9F%E8%AE%A1%E6%8A%A5%E5%91%8A/%E8%B7%A8%E6%80%A7%E5%88%AB%E6%A0%A1%E5%9B%AD%E4%B8%8E%E8%AF%81%E4%BB%B6%E4%BF%AE%E6%94%B9%E8%B0%83%E6%9F%A5%E6%8A%A5%E5%91%8A_page/
- Chinese Transgender Digital Archive. 2026c. “跨性别证件修改手册.” https://digital.transchinese.org/%E7%A4%BE%E7%BE%A4%E5%8F%8ANGO%E6%96%87%E4%BB%B6/%E6%89%8B%E5%86%8C%E6%8C%87%E5%8D%97/%E8%B7%A8%E6%80%A7%E5%88%AB%E8%AF%81%E4%BB%B6%E4%BF%AE%E6%94%B9%E6%89%8B%E5%86%8C_page/
- Amnesty International. 2019a. China: “I Need My Parents’ Consent to Be Myself”: Barriers to Gender-Affirming Treatments for Transgender People in China. https://www.amnesty.org/en/documents/asa17/0269/2019/en/
- Amnesty International. 2019b. “我需要家长同意才能做自己——中国跨性别者寻求性别确认医疗程序时遇到的障碍.” https://www.amnesty.org/zh-hans/documents/asa17/0269/2019/zh-hans/
- Amnesty International. 2019c. Report PDF. https://www.amnesty.org/en/wp-content/uploads/2021/05/ASA1702692019ENGLISH.pdf
- CNLGBTDATA. 2023a. “2021全国跨性别健康调研报告.” https://cnlgbtdata.com/doc/314/
- Chinese Transgender Digital Archive. 2026d. “2021全国跨性别健康调研报告.” https://digital.transchinese.org/%E7%A4%BE%E7%BE%A4%E5%8F%8ANGO%E6%96%87%E4%BB%B6/%E7%BB%9F%E8%AE%A1%E6%8A%A5%E5%91%8A/2021%E5%85%A8%E5%9B%BD%E8%B7%A8%E6%80%A7%E5%88%AB%E5%81%A5%E5%BA%B7%E8%B0%83%E7%A0%94%E6%8A%A5%E5%91%8A_page/
- PUH3 Department of Plastic Surgery. 2018. “北医三院易性症序列医疗团队成立.” https://www.sar.com.cn/xinwen/news/12295.html
- Common Language. 2020. “影响性诉讼⑦:这次,把‘性别认同及性别表达’写进判决.” https://commonlanguage.github.io/TYarchives2020/20200518_1%E5%BD%B1%E5%93%8D%E6%80%A7%E8%AF%89%E8%AE%BC%E2%91%A6%E8%BF%99%E6%AC%A1%EF%BC%8C%E6%8A%8A%E2%80%9C%E6%80%A7%E5%88%AB%E8%AE%A4%E5%90%8C%E5%8F%8A%E6%80%A7%E5%88%AB%E8%A1%A8%E8%BE%BE%E2%80%9D%E5%86%99%E8%BF%9B%E5%88%A4%E5%86%B3.html
- Chinese Transgender Digital Archive. 2026e. “Technical Management Specification for Sex Reassignment Surgery (Trial).” https://digital.transchinese.org/%E6%94%BF%E5%BA%9C%E5%8F%8A%E5%AE%98%E6%96%B9%E7%BB%84%E7%BB%87%E6%96%87%E4%BB%B6/%E4%B8%AD%E5%9B%BD%E5%A4%A7%E9%99%86/%E5%8F%98%E6%80%A7%E6%89%8B%E6%9C%AF%E6%8A%80%E6%9C%AF%E7%AE%A1%E7%90%86%E8%A7%84%E8%8C%83%EF%BC%88%E8%AF%95%E8%A1%8C%EF%BC%89/
- Chinese Transgender Digital Archive. 2026f. Homepage. https://digital.transchinese.org/
- Common Crawl. 2026. “Open Repository of Web Crawl Data.” https://commoncrawl.org/
- Memento. 2026. “Time Travel for the Web.” https://mementoweb.org/about/