<?xml version="1.0" encoding="UTF-8"?><TEI xmlns="http://www.tei-c.org/ns/1.0" rend="tei2017" xml:id="t_100_hannesschlager_andorfer_gender">
  <teiHeader>
      <fileDesc>
         <titleStmt>
            <title type="main">Sex in the TEI: The TEI 2016 gender check</title>
            <author>
               <name n="hannesschläger_v">
                  <forename>Vanessa</forename> 
                  <surname>Hannesschläger</surname>
               </name>
               <affiliation>Vanessa Hannesschläger is a researcher at the Austrian Centre for Digital Humanities of the Austrian Academy of Sciences (ACDH-OEAW), where she is responsible for legal issues. She is involved in several projects in which she works on data modelling, digital editing, and in the outreach department. In addition, she is completing her PhD with the German department of the University of Vienna. Her research interests include legal frameworks of digital research, biography theory, archive theory, modern Austrian literature, and the contemporary developments of gender issues in society. For more information, please visit <ref target="http://vanessahannesschlaeger.wordpress.com/">http://vanessahannesschlaeger.wordpress.com/</ref> .</affiliation>
               <email>vanessa.hannesschlaeger@oeaw.ac.at</email>
            </author>
            <author>
               <name n="andorfer_p">
                  <forename>Peter</forename>
                  <surname>Andorfer</surname>
               </name>
               <affiliation>Peter Andorfer studied history at Innsbruck University, where he finished a PhD in history with a thesis on the works of the Tyrolean peasant <ref target="https://github.com/csae8092/MWB">Leonhard Millinger</ref> (1753–1834). During an extended research period at the Herzog August Bibliothek in Wolfenbüttel (Lower Saxony, Germany), financed by a <soCalled>Digital-Humanities Scholarship</soCalled>, he published an online edition of Millinger’s main work <ref target="http://diglib.hab.de/edoc/ed000223/start.htm">The Depiction of the World</ref>. He has also worked on the topics <soCalled>research data</soCalled> and <soCalled>scientific collections</soCalled> in <ref target="https://de.dariah.eu/">DARIAH-DE</ref> and maintains the webpage <ref target="http://www.digital-archiv.at/">www.digital-archiv.at</ref> for developing and deploying different kinds of DH-projects.</affiliation>
               <email>peter.andorfer@oeaw.ac.at</email>
            </author>
         </titleStmt>
         <publicationStmt>
            <publisher>TEI Consortium</publisher>
            <date/>
            <availability>
               <p>Creative Commons Attribution 4.0 International</p>
            </availability>
         </publicationStmt>
         <sourceDesc>
            <p>No source, born digital.</p>
         </sourceDesc>
      </fileDesc>
      <encodingDesc>
         <projectDesc>
            <p>TEI 2017 Conference Abstracts.</p>
         </projectDesc>
      </encodingDesc>
      <profileDesc>
         <langUsage>
            <language ident="en">en</language>
         </langUsage>
         <textClass>
            <classCode scheme="conference">paper</classCode>
            <keywords xml:lang="en">
               <term>gender</term>
               <term>conference contributors</term>
               <term>data enrichment</term>
            </keywords>
         </textClass>
      </profileDesc>
      <revisionDesc>
         <change who="TEH" when="2017-08-21">Tracey El Hajj encoded the file</change>
      </revisionDesc>
  </teiHeader>
  <text>
      <front>
         <div type="abstract" xml:id="abstract"><p>We decided to genderize the <gi>forename</gi>s rather
            than the <gi>person</gi>s and will explain this decision during our talk with reference to
               contemporary gender theory.</p></div>
      </front>
      <body>
         <p>The abstracts of the TEI Conference and Members’ Meeting 2016 were published by the
            hosts (Austrian Centre for Digital Humanities / Austrian Academy of Sciences) as TEI
            encoded XML documents on GitHub (<ref type="bibl" target="#HannesschlagerSchopper2016">Hannesschläger and Schopper 2016</ref>). This, and the fact that these documents were published under a CC-BY-SA-4.0 license, made it possible to take these data and <soCalled>play</soCalled> with
            them - for instance by building a web application to publish as well as analyse the data.</p>
         <p>Among other things, the editors tagged the forenames of the authors with the according
            <gi>forename</gi>. This allowed us to ask the question about gender distribution among the
            contributors to the conference. What started as a playful exercise in data mining, processing,
            and analysis, lead to categorical questions about how to assign and especially how to
            encode gender information to persons. We decided to genderize the <gi>forename</gi>s rather
               than the <gi>person</gi>s and will explain this dencision during our talk with reference to
                  contemporary gender theory.</p>
         <p>As far as alignment of forenames and gender is concerned, this is a simple task, at least
            from a technical point of view. As described on the tei2016app website in detail (<ref type="bibl" target="#AndorferHannesschlager2016">Andorfer and Hannesschläger 2016</ref>)
            looked for a comprehensive and structured list of forenames that have already been mapped
            to genders, e.g., a list of female forenames and a list of male forenames. Secondly, the
            tagged <gi>forename</gi> of the respective TEI abstract had to be checked against these lists.</p>
         <p>The first list of gendered names we found was is the one provided by Mark Kantrowitz that is
            used e.g., in the NLTK package (<ref type="bibl" target="#Kantrowitz2017">Kantrowitz 2017</ref>). 
            This list was ingested into a django-based web service and
            accessed by an XQuery script, iterating through all forename elements of the abstracts
            corpus, sending each forename to the service’s endpoint and storing the returned answer.</p>
         <p>While simple from a technical viewpoint, from a gender studies viewpoint this approach was
            questionable because Kantrowitz does not provide information on how the list was compiled
            or what criteria were applied to group names into the categories <mentioned>female</mentioned>, <mentioned>male</mentioned>, and <mentioned>pet</mentioned>.</p>
         <p>Other sources like e.g., <ref target="http://genderize.io">genderize.io</ref> do not only provide more data, but also give information
            about how the data was gathered and categorized. The most important argument for this
            data source was <ref target="http://genderize.io">genderize.io</ref>’s claim that the data collected there was assembled by scraping data from social network profiles, where people can declare their gender
            themselves. 
         <note>
            <p>However, it has to be mentioned that we do not have full confidence in the truth of the claim that the data
               was gathered from social networks because <ref target="http://genderize.io">genderize.io</ref> only knows two genders, but platforms like Facebook already offer many more choices.</p>
         </note></p>
         <p>Solving the issue of finding an adequate data source led to the question of how to encode
            this scraped information in a useful and TEI conformant way. The <gi>sex</gi> tag only allows to
               encode assumptions about a person’s sex, and <gi>gender</gi> about morphological gender of a
                  lexical item, but neither of this fits our needs as we wanted to encode the gender a forename
                  is most commonly associated with. As it turned out, the broader issue of how to encode a
                  person’s sex has lead to quite some lengthy debates in the TEI community,<note>
                     <p>E.g., <ptr target="https://github.com/TEIC/TEI/issues/426"/></p>
                  </note> non of which
            consider the distinction between sex and gender (<ref type="bibl" target="#WestZimmerman1987">West and Zimmerman 1987</ref>)
                  or discuss the questionable praxis of
                  assigning either to a person other than oneself. The discussions focus on which values
                  should be used (allowed) to encode a person’s sex but do not consider the question on if
                  and how a forename element could/should be gendered.</p>
         <p>For the current project, we <soCalled>solved</soCalled> this issue by encoding the <gi>forename</gi>’s gender with the
            help of a <att>type</att>. Concerning the values of these attributes, we came across the same issues
            that were discussed in context of <gi>sex</gi>, e.g., Should we encode gender information following
               some (iso)standard or choose custom/arbitrary values? Finally, we decided to chose the
               values <val>female</val>, <val>male</val>, and <val>nomatch</val> (the latter for forenames that did not match any name
            gendered by <ref target="http://genderize.io">genderize.io</ref> - and <ref target="http://genderchecker.com">genderchecker.com</ref>, which was used to reconcile names not
            found in <ref target="http://genderize.io">genderize.io</ref>).</p>
         <p>As a result, we can now say that 87 forenames of the contributors of the TEI-conference
            were male, 37 female and three <val>no-matches</val>. Concerning authorship of the published
            papers, there were 39 abstracts with more male than female author’s forenames, in 14
            abstracts more female than male, eight texts with an equal distribution and two abstracts
            with an unclear result (meaning that most names couldn’t clearly be assigned a male or
            female gender).</p>
      </body>
    
      <back>
         <div type="bibliography">
            <listBibl>
               <bibl xml:id="AndorferHannesschlager2016"><author>Andorfer, Peter</author>, and <author>Vanessa Hannesschläger</author>. <date>2016</date>. <title level="a">Gender distribution among the contributors to TEI 2016.</title> <title level="j">tei2016app</title>. <ref target="https://tei2016app.acdh.oeaw.ac.at/pages/show.html?document=genderize.xml&amp;directory=meta&amp;stylesheet=meta"/>.</bibl>
               <bibl xml:id="HannesschlagerSchopper2016"><editor>Hannesschläger, Vanessa</editor>, and <editor>Schopper, Daniel</editor> <date>2017</date>. <title level="a">Book of Abstracts in TEI XML</title>. TEI Conference and Members’ Meeting 2016. <ref target="https://github.com/acdh-oeaw/TEI2016abstracts"/>.</bibl>
               <bibl xml:id="Kantrowitz2017"><author>Kantrowitz, Mark</author>. <date>2017</date>. <title level="a">Name Corpus: List of Male, Female, and Pet names.</title> <title level="j">CMU Artificial Intelligence
                  Repository.</title> Last
                  modified:
                  02-Apr-1997. <ref target="http://www.cs.cmu.edu/afs/cs/project/ai-repository/ai/areas/nlp/corpora/names/"/></bibl>
               <bibl xml:id="WestZimmerman1987"><author>West, Candace</author>, and <author>Don H. Zimmerman.</author><date>1987</date>. <title level="a">Doing Gender.</title> <title level="j">Gender and Society.</title> <biblScope unit="volume">1</biblScope>(<biblScope unit="issue">2</biblScope>): <biblScope unit="page">125–151</biblScope>.</bibl>
            </listBibl>
         </div>
      </back>
  </text>
</TEI>