Image from Huffington Post
While America’s 58th presidential election divided the nation’s population, you’d be hard-pressed to find a single reputable poll predicting a win for the Republican nominee, Donald Trump, in late 2016.
Battling PR blunder after blunder, the Trump campaign seemed to spiral out of control as election day drew nearer. Each week brought a new scandal, alienating a new demographic. That’s not to say the Clinton campaign was without its dark spots, however. Trump dug his heels in on growing concerns over an FBI investigation regarding Hillary Clinton’s private email server, and even called into question ex-president and hopeful “First Gentleman”, Bill Clinton.
Despite coming under fire, the outlook was bright for the Clinton campaign. In fact, even the most pessimistic models gave democratic nominee Hillary Clinton a ~70% chance of winning the election:
From statistics celeb Nate Silver’s Fivethirtyeight, to the New York Times and LA Times, the list of outlets predicting a Clinton win was endless.
When November 8th rolled by, the nation watched as one of the biggest upsets in the history of presidential elections unfolded before their eyes. Donald Trump defied predictions as he picked up electoral votes at a breakneck pace.
Clinton: Blue, Trump: Red
Final tally of electoral votes: Trump 306 - Clinton 232
Now, months after his victory and knee deep into the Trump administration, this analysis aims to investigate key characteristics of voters in this election, with the goal of better understanding why Clinton lost the 2016 election.
A common storyline after the election was that Hillary lost the “middle class” vote. Historically, the Democratic party dominates the vote among lower income individuals. This tendency shifted during the 2016 election, as shown here among swing states:
The vertical line labeled 67,406 represents the average median income in all states
There is a strong trend between income and candidate choice in the swing states. 4 out of 5 of Hillary’s victories were in states whose median incomes push well past the country average median income of $67,406, while all of Trump’s states are below or barely reach the average line.
Almost all of the swing states were highly contentious (“Narrow Wins”), with only Iowa and Ohio being won comfortably by Trump (read more about victory margin classifications here in the notes about the data).
However, Trump did win in more swing states than Hillary, bagging 7 out of 12 states, which contributed in part to his significant electoral vote lead.
Once again the states won by Hillary have noticably higher average median incomes across the board.
Hillary seemed to do well in more educated states. Comparing states by bachelor’s degree attainment, Trump dominated regions with relatively low attainment rates:
This correlation is particularly visible in the South and Midwest, both regions with low degree attainment rates and a high amount of Trump states. The highly educated Northwest was primarily democratic, and the West generally had a variety of both education rates and election winners.
Extending our view to other education levels, the trend continues.
States with higher percentages of people without high school educations had less democratic votes, although with a weak correlation.
States with higher bachelor’s degree rates had a stronger correlation with more Democratic votes.
Graduate degree rates were the best predictor for Democratic vote percentages, with the highest R-squared value (.526) of the three education categories.
Given the candidates’ disparate views on healthcare, it would be a reasonable assumption that uninsured voters would tend to support Clinton, who promised to uphold Obamacare. Surprisingly, this is not true:
States with a higher uninsured population had less Democratic votes. However, when factoring in poverty in these states, we see a familiar trend emerge.
States with high amounts of poverty also tended to be uninsured, and poorer states tended to vote Republican in this election, as discussed in the “Income” section above.
We can confirm this relationship by looking at uninsurance and income:
As expected, lower incomes generally mean less insured people. Despite Hillary’s more progressive insurance plan, low income individuals did not for her.
Most minority groups are consistently spread out with respect to Democratic voting percentages.
This is consistent with reports that Clinton was unable to rally minority voters.
While certain demographics may not have shown up at the polls, voter engagement was fairly even between states, regardless of the winning candidate.
Note: Voting Ratio is calculated based on total population of a state, not by registered voters, as “voter turnout” is usually calculated.
The Clinton campaign lost the vote of the “average American”, while winning among more highly educated and wealthier populations (although our analysis did not touch on the voting tendencies of the “ultra rich”).
Despite having policies that were arguably better for poorer individuals (such as in the case of health care), Hillary lost many votes from this traditionally Democratic voting block, who showed up in record numbers for Obama in 2008 but were uninspired by the Clinton campaign.
Should Clinton run against Trump again in 2020, the key takeaways for her campaign to regain the working class vote are:
1. A more focused economic message
As a part of a bigger focus on social equity and diversity, Clinton’s campaign stressed wealth redistribution and increased benefits.
“Despite what you hear, we don’t need to make America great again. America has never stopped being great. But we do need to make America whole again. Instead of building walls, we need to be tearing down barriers,” Clinton said. Source
While these goals resonated with her already firm base of the educated and financially secure, an emphasis on job creation may have captured the attention of the lower/lower-middle class more strongly.
Trump’s campaign promises of manufacturing jobs and protectionism was much more convicing to those worried about economic and job outlook.
After all, if you’ve been worried about keeping your job and paying your bills, would you rather hear Trump say he’ll “Make America Great Again”, or hear Hillary say “we don’t need to make America great again”?
2. Image control
While both candidates were wealthy, Clinton allowed herself to be painted as the candidate representing wealthy interests.
While she struck back by calling out Trump’s inexperience and bigotry in the debates, Clinton rarely contested the idea that she was out of touch with the working class. Had she attacked Trump on being born into extreme wealth and being handed millions of dollars, she might have regained some of her image in the eyes of the working class.
All relevant data is hosted publically here on data.world as well as here on Github.
The primary dataset was aggregated from county level info, so all ratios are averages of county percentages. For example, Texas is listed as having an uninsured rate of 26.6%. Here, this means that the “average” or typical Texas county has 26.6% people uninsured, NOT that 26.6% of all Texans are uninsured. This caveat is not true of the democratic/republican vote percentages, which were taken directly from the Politico elections result report.
Some metrics used frequently in our analysis do not appear in our dataset. The most prominent of these metrics is “Victory Margin”, which is the margin a candidate won a state by, divided into categories. The cutoffs are arbitrary, but help to see how many states were closely competitive, mildly competitive, or not competitive at all. The calculation is shown in the following formula:
Other calculated metrics that are not in our dataset, such as Poverty Ratio and Voter Ratio, are simple normalizations so that we may compare across states with disparate populations. For example, Poverty Ratio is the number of people in poverty in a state divided by the total population of that state.
The remaining data was collected by the New York Times and The Ulster Institute for Social Research for a publication in the Open Quantitative Sociology & Political Science journal in 2016. Also included are extracts from the 2015 Census Data datasets for income, poverty, and race that were joined to the elections data set. The sources for these extracts can be found on the U.S. Census Bureau’s data.world profile.
The interactive visualizations can be accessed here, with instructions for the interactions at the top of each page.
The following query is used by the Shiny application to generate the web based visualizations. The query was passed into the data.world R package, which has slightly different syntax than standard SQL.
select E.*, I.Median_Income, P.Number_Below_PovertyLine, R.White_Population from `1ElectionsData.csv/1ElectionsData` E left join `acs-2015-5-e-income-medians.csv/acs-2015-5-e-income-medians` I on (E.State = I.AreaName) left join `acs-2015-5-e-poverty-populationinpoverty.csv/acs-2015-5-e-poverty-populationinpoverty` P on (E.State = P.AreaName) left join `acs-2015-5-e-race-whitepopulation.csv/acs-2015-5-e-race-whitepopulation` R on (E.State = R.AreaName) order by E.State
Census data often comes in confusing formats, with coded column names and many tables. All census data used here was taken from the table of state level data, referred to as USA_All_States within each census dataset. The description for the coded column name will be provided along with each query, but have been edited for readability.
The column name B19013_001 has the description “Median household income in the past 12 months (in 2015 Inflation-adjusted dollars).” The dataset that was queried can be found here.
select AreaName, B19013_001 as Median_Income from USA_All_States
The column name B17001_002 has the description “Population For Whom Poverty Status Is Determined With an Income in the past 12 months below poverty level.” The dataset that was queried can be found here.
select AreaName, B17001_002 as Number_Below_PovertyLine from USA_All_States
The column name B17001_002 has the description “Race for Total Population (White alone).” The dataset that was queried can be found here.
select AreaName, B17001_002 as White_Population from USA_All_States
While the full dataset is downloadable as linked above, summary statistics of our data (after joins) are provided here.
## State Region Total.Population electoralvotes ## Length:51 Midwest :12 Min. : 543788 Min. : 3.00 ## Class :character Northeast:10 1st Qu.: 1664655 1st Qu.: 4.50 ## Mode :character South :16 Median : 4295684 Median : 8.00 ## West :13 Mean : 5982422 Mean :10.55 ## 3rd Qu.: 6562275 3rd Qu.:11.50 ## Max. :36781242 Max. :55.00 ## ## rep16_frac dem16_frac votes votes16_trumpd ## Min. :0.0410 Min. :0.2250 Min. : 248742 Min. : 11553 ## 1st Qu.:0.4150 1st Qu.:0.3610 1st Qu.: 736890 1st Qu.: 376494 ## Median :0.4910 Median :0.4670 Median : 1923346 Median : 947934 ## Mean :0.4912 Mean :0.4501 Mean : 2552568 Mean :1199907 ## 3rd Qu.:0.5765 3rd Qu.:0.5255 3rd Qu.: 3094736 3rd Qu.:1545866 ## Max. :0.7010 Max. :0.9280 Max. :11954317 Max. :4681590 ## ## votes16_clintonh votes16_johnsong votes16_steinj ## Min. : 55949 Min. : 4501 Min. : 2512 ## 1st Qu.: 274023 1st Qu.: 28887 1st Qu.: 8000 ## Median : 779535 Median : 57322 Median : 14075 ## Mean :1225916 Mean : 83822 Mean : 29245 ## 3rd Qu.:1723912 3rd Qu.:125718 3rd Qu.: 36957 ## Max. :7362490 Max. :402406 Max. :220312 ## NA's :6 ## At.Least.Bachelor.s.Degree At.Least.High.School.Diploma ## Min. :13.79 Min. :43.25 ## 1st Qu.:17.42 1st Qu.:80.65 ## Median :20.02 Median :86.17 ## Mean :21.31 Mean :83.76 ## 3rd Qu.:23.76 3rd Qu.:87.96 ## Max. :36.41 Max. :91.15 ## ## Less.Than.High.School Graduate.Degree White.not.Latino.Population ## Min. : 6.75 Min. : 4.298 Min. :28.68 ## 1st Qu.:11.41 1st Qu.: 5.547 1st Qu.:67.81 ## Median :13.66 Median : 6.435 Median :82.53 ## Mean :15.05 Mean : 7.592 Mean :77.17 ## 3rd Qu.:18.27 3rd Qu.: 8.554 3rd Qu.:90.36 ## Max. :24.26 Max. :15.300 Max. :95.48 ## ## African.American.Population Native.American.Population ## Min. : 0.2509 Min. : 0.1463 ## 1st Qu.: 1.0101 1st Qu.: 0.2740 ## Median : 3.1319 Median : 0.5300 ## Mean : 8.2964 Mean : 2.4260 ## 3rd Qu.: 9.3464 3rd Qu.: 1.7135 ## Max. :52.3000 Max. :31.6190 ## ## Asian.American.Population Population.some.other.race.or.races ## Min. : 0.3545 Min. : 0.7713 ## 1st Qu.: 0.5965 1st Qu.: 1.2077 ## Median : 0.8464 Median : 1.5454 ## Mean : 1.9175 Mean : 2.4377 ## 3rd Qu.: 1.7243 3rd Qu.: 1.9259 ## Max. :27.3700 Max. :34.6600 ## ## Latino.Population Management.professional.and.related.occupations ## Min. : 0.9164 Min. :25.99 ## 1st Qu.: 2.7542 1st Qu.:28.28 ## Median : 4.3708 Median :30.79 ## Mean : 7.7546 Mean :31.69 ## 3rd Qu.: 8.3650 3rd Qu.:33.94 ## Max. :45.2500 Max. :56.75 ## ## Service.occupations Sales.and.office.occupations ## Min. :15.36 Min. :19.20 ## 1st Qu.:16.81 1st Qu.:22.40 ## Median :17.22 Median :23.11 ## Mean :17.64 Mean :23.05 ## 3rd Qu.:18.30 3rd Qu.:23.96 ## Max. :22.82 Max. :26.29 ## ## Farming.fishing.and.forestry.occupations ## Min. :0.1000 ## 1st Qu.:0.9921 ## Median :1.7618 ## Mean :1.9235 ## 3rd Qu.:2.4727 ## Max. :5.3977 ## ## Construction.extraction.maintenance.and.repair.occupations ## Min. : 3.300 ## 1st Qu.: 9.966 ## Median :10.804 ## Mean :11.181 ## 3rd Qu.:12.183 ## Max. :16.706 ## ## Production.transportation.and.material.moving.occupations ## Min. : 4.65 ## 1st Qu.:11.10 ## Median :14.07 ## Mean :14.52 ## 3rd Qu.:18.24 ## Max. :22.80 ## ## Adult.obesity Diabetes Uninsured Unemployment ## Min. :0.2068 Min. :0.06306 Min. :0.0535 Min. :0.03573 ## 1st Qu.:0.2662 1st Qu.:0.08804 1st Qu.:0.1249 1st Qu.:0.06717 ## Median :0.2981 Median :0.09685 Median :0.1741 Median :0.08051 ## Mean :0.2921 Mean :0.10115 Mean :0.1664 Mean :0.07832 ## 3rd Qu.:0.3185 3rd Qu.:0.11154 3rd Qu.:0.2018 3rd Qu.:0.09262 ## Max. :0.3670 Max. :0.14082 Max. :0.2760 Max. :0.12252 ## ## White_Population Number_Below_PovertyLine Median_Income ## Min. : 260325 Min. : 64995 Min. :49274 ## 1st Qu.: 1503912 1st Qu.: 238146 1st Qu.:58720 ## Median : 3210708 Median : 636947 Median :66389 ## Mean : 4567511 Mean : 936256 Mean :67405 ## 3rd Qu.: 5481689 3rd Qu.: 961445 3rd Qu.:74035 ## Max. :23747013 Max. :6135142 Max. :90089 ##
The following tools were used:
Python (to clean the dataset using the “Pandas” library)
SQL (to query census data and aggregate our state level data from a county level dataset)
Tableau (to create the visualizations used on this page)
R (to create interactive web visualizations using “ggplot” and "Shiny", and the "Knitr" package to generate this page)
data.world and Github (to host and share the dataset)