Data Analysis Case Study

Objectives:
The purpose of this exam is to help you synthesize and make further connections between and across the readings. You need to make an informed argument supported by what you’ve learned this semester. Demonstrate your mastery of the materials we read.
Rules:
● Answer 1 from the descriptive section and answer 1 from the case study section.
● Each answer should be at least 700 words and should be no more than 1400 words. Your answers are complete when you feel you have articulated a thorough and complete argument.
● Use and reference as many of the books and journal articles that you read this semester as relevant to help you make your argument.
● Be sure to demonstrate mastery of the books and two journal articles (data journalism and self-tracking) you read this semester and answer the prompt fully.
● Be sure to use concepts and define key terms from our authors (and cite whose definitions and concepts you’re using).
● Only cite from the given readings.
● Lectures, discussion on Blackboard, and the writing projects all may be used to help you support and justify your answers.

Readings:
Neff & Nafus, Self-Tracking
Self-Tracking, Chapter 2
Self-Tracking, Chapter 3
Self-Tracking, Chapter 5
Self-Tracking and the issue of agency:
• Crawford, K., Lingel, J., & Karppi, T. (2015). Our metrics, ourselves: A hundred years of self-tracking from the weight scale to the wrist wearable device. European Journal of Cultural Studies, 18(4-5), 479-496. (Available as PDF)

D’Ignazio & Klein, Data Feminism (Videos)
Data Feminism, Chapter 1 – “The Power Chapter”

Data Feminism, Chapter 4 – “What Gets Counted Counts”

Data Feminism, Chapter 6 – “The Numbers Don’t Speak for Themselves”

Data Feminism, Chapter 7 – “Show your work”

John Cheney-Lippold, We Are Data
We Are Data, “Introduction,” p. 3- 33
We are Data, Chapter 1, “Categorization: Making Data Useful,” p. 39-55
We are Data, Chapter 1, “Categorization: Making Data Useful,” p. 55-72
We are Data, Chapter 1, “Categorization: Making Data Useful,” p. 73-92
We are Data, Chapter 4, “Privacy: Wanted Dead or Alive”

Parasie-Dagiral2012_Article_DataDrivenJournalism

Descriptive Prompts:
1) The central challenge we have discussed in this class is the issue of use of data and algorithms for creating systems that are becoming integral to our day-to-day existence. Be it social media, recommendation systems (such as Amazon, Netflix) or even self-tracking technology such as Fitbits. Although, these systems are in many ways enabling and give us the ability to connect, share and express ourselves in many different ways, they are also in many ways constraining, as we have discussed through the many examples in class. Particularly, the question(s) that arise are: Are these systems designed for all? If not, whom do they miss out and why? Where do designers fall short? Using concepts, examples in the text, write a response to the above question and to this prompt. Your answer must depict arguments from at least two of the four books we have studied.

[Hint: Think about the idea of ‘systems of exclusion’ and connect it with concepts such as algorithmic citizenship, think about privilege hazard, technochauvenism and how that creates issues of marginalization, think about the issue of bias and assumptions ingrained into these systems]

Case study prompts:
A premier data science focused organization has now decided that they wish to procure and analyze the social media data generated by employees to better understand their employees and also potentially use insights from this data to form teams for collaborative projects across departments. The goal and the vision of the organization is that in addition to using professional footprints, using such information will better align people’s interests and thereby increase overall productivity. To create this system, a survey was distributed across the employees of the organization to collect data on various personal attributes – the survey had questions related to demographics, personal interests and certain psychological markers that were derived from prior literature on sports teams. Employees were also required to share their Facebook handles if they had one. Based on this survey a team was selected comprising of data scientists to analyze this data and devise a strategy to find the best possible matches. The machine learning driven recommendations were crafted using unsupervised learning techniques we have studied in class. No tests were performed on the results and directly these were implemented through a trial for a three-month period. Although still used, employees began to express their discontentment with such a system. You goal now is to analyze where potentially the system and those envisioning such technocentric approaches fell short.

Based on this scenario, use concepts from Data Feminism, Artificial unintelligence and We are Data to articulate your response to the following questions:
a. Do you agree with this vision? What are potential ethical challenges that may arise in this scenario?
b. Think about the data collected, analysis process and the overall AI pipeline, where do you think the system developers and planners fall short?
c. Think about the large-scale implications of using such systems on a day to day basis, using concepts from the class, ideate on potential negative ramifications and challenges introduced by extended appropriation of such systems.