Who Says What to Whom on Twitter

Shaomei Wu∗
Cornell University, USA

Jake M. Hofman hofman@yahoo-inc.com Duncan J. Watts
Yahoo! Research, NY, USA Yahoo! Research, NY, USA

sw475@cornell.edu Winter A. Mason
Yahoo! Research, NY, USA

winteram@yahooinc.com

djw@yahoo-inc.com

ABSTRACT
We study several longstanding questions in media communications research, in the context of the microblogging service Twitter, regarding the production, flow, and consumption of information. To do so, we exploit a recently introduced feature of Twitter—known as Twitter lists—to distinguish between elite users, by which we mean specifically celebrities, bloggers, and representatives of media outlets and other formal organizations, and ordinary users. Based on this classification, we find a striking concentration of attention on Twitter—roughly 50% of tweets consumed are generated by just 20K elite users—where the media produces the most information, but celebrities are the most followed. We also find significant homophily within categories: celebrities listen to celebrities, while bloggers listen to bloggers etc; however, bloggers in general rebroadcast more information than the other categories. Next we re-examine the classical “two-step flow” theory of communications, finding considerable support for it on Twitter, but also some interesting differences. Third, we find that URLs broadcast by different categories of users or containing different types of content exhibit systematically different lifespans. And finally, we examine the attention paid by the different user categories to different news topics.

1.

INTRODUCTION

Categories and Subject Descriptors
H.1.2 [Models and Principles]: User/Machine Systems; J.4 [Social and Behavioral Sciences]: Sociology

General Terms
two-step flow, communications, classification

Keywords
Communication networks, Twitter, information flow
∗ Part of this research was performed while the author was visiting Yahoo! Research, New York.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. To copy otherwise, to republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. WWW ’11 Hyderabad, India Copyright 20XX ACM X-XXXXX-XX-X/XX/XX ...$10.00.

A longstanding objective of media communications research is encapsulated by what is known as Lasswell’s maxim: “who says what to whom in what channel with what effect” [9], so-named for one of the pioneers of the field, Harold Lasswell. Although simple to state, Laswell’s maxim has proven difficult to satisfy in the more-than 60 years since he stated it, in part because it is generally difficult to observe information flows in large populations, and in part because different channels have very different attributes and effects. As a result, theories of communications have tended to focus either on “mass” communication, defined as “oneway message transmissions from one source to a large, relatively undifferentiated and anonymous audience,” or on “interpersonal” communication, meaning a “two-way message exchange between two or more individuals.” [13]. Correspondingly, debates among communication theorists have tended to revolve around the relative importance of these two putative modes of communication. For example, whereas early theories such as the so-called “hypodermic model” posited that mass media exerted direct and relatively strong effects on public opinion, mid-century researchers [10, 6, 11, 4] argued that the mass media influenced the public only indirectly, via what they called a two-step flow of communications, where the critical intermediate layer was occupied by a category of media-savvy individuals called opinion leaders. The resulting “limited effects” paradigm was then subsequently challenged by a new generation of researchers [5], who claimed that the real importance of the mass media lay in its ability to set the agenda of public discourse. But in recent years rising public skepticism of mass media, along with changes in media and communication technology, have tilted conventional academic wisdom once more in favor of interpersonal communication, which some identify as a “new era” of minimal effects [2]. Recent changes in technology, however, have increasingly undermined the validity of the mass vs. interpersonal dichotomy itself. On the one hand, over the past few decades mass communication has experienced a proliferation of new channels, including cable television, satellite radio, specialist book and magazine publishers, and of course an array of web-based media such as sponsored blogs, online communities, and social news sites. Correspondingly, the traditional mass audience once associated with, say, network television has fragmented into many smaller audiences, each of which increasingly selects the information to which it is exposed, and in some cases generates the information itself. Mean-

kaist. ranging from less than a day to months. Twitter is primarily made up of many millions of users who seem to be ordinary individuals communicating with their friends and acquaintances in a manner largely consistent with traditional notions of interpersonal communication. The remainder of the paper proceeds as follows. and that different content types exhibit dramatically different characteristic lifespans. however. Therefore in this paper we limit our focus to the “who says what to whom” part of Laswell’s maxim. often managed by themselves or publicists. [8]. especially as Twitter—unlike television. In section 4. a number of recent papers have examined information diffusion on Twitter. a new class of “semi-public” individuals like bloggers. and social networking sites to afford individuals ever-larger audiences. and second the lifespan of URLs as a function of their origin and their content. and show that they deliver qualitatively similar results. This data represents a crawl of the graph seeded with all users on Twitter as observed by July 31st. Nowhere is the erosion of traditional categories more apparent than in the micro-blogging platform Twitter. celebrities. Next.1 DATA AND METHODS Twitter Follower Graph In order to understand how information is transmitted on Twitter. we need to know the channels by which it flows. which included 42M users and 1. In section 5. finding that although users with large follower counts and past success in triggering cascades were on average more likely to trigger large cascades in the future. In addition. and is publicly available1 .1. page-rank. we revisit the theory of the two-step flow—arguably the dominant theory of communications for much of the past 50 years—finding considerable support for the theory as well as some interesting differences. further classifying elite users into one of four categories of interest— media. Together. Twitter. such as changes in behavior. • We find that different categories of users place slightly different emphasis on different types of content. and number of retweets—finding that the ranking of the most influential users differed depending on the measure. and print media—enables one to easily observe information flows among the members of its ecosystem. in addition to conventional celebrities. In the next section. finding that although audience attention is highly concentrated on a minority of elite users. Our paper differs from this earlier work by shifting attention from the ranking of individual users in terms of various influence measures to the flow of information among different categories of users.. in section 6 we conclude with a brief discussion of future work. examining first who shares what kinds of media content. As reported by Kwak et al. Bakshy et al. much of the information they produce reaches the masses indirectly via a large population of intermediaries. To illustrate. Third. thus bypassing the traditional intermediation of the mass media between celebrities and fans. 2. media organizations—along with corporations. we review related work. these two trends have greatly obscured the historical distinction between mass and interpersonal communications. RELATED WORK Aside from the communications literature surveyed above.3 in which we describe how we use Twitter Lists to classify users. our paper makes three main contributions: • We introduce a method for classifying users using Twitter Lists into “elite” and “ordinary” users. we consider “who listens to what”. [8] studied the topological features of the Twitter follower graph. [1] studied the distribution of retweet cascades on Twitter. 3. therefore. concluding from the highly skewed nature of the distribution of followers and the low rate of reciprocated ties that Twitter more closely resembled an information sharing network than a social network— a conclusion that is consistent with our own view. In particular. In section 4 we analyze the production of information on Twitter. and bloggers. Finally. 2009.kr/traces/WWW2010. these individuals communicate directly with their millions of followers. again finding that ranking depended on the influence measure. who is following whom on Twitter. leading some scholars to refer instead to “masspersonal” communications [13]. In section 3 we discuss our data and methods. in some cases becoming more prominent than traditional public figures such as entertainers and elected officials. Weng et al. Unfortunately. attitudes. To this end. • We investigate the flow of information among these categories. etc.html from . email lists. the follower graph is a directed network characterized by highly skewed distributions both 1 The data is free to download http://an. 3. compared three different measures of influence—number of followers. In a similar vein. and subject matter experts have come to occupy an important niche on Twitter. as well as how information originating from traditional media sources reaches the masses. we are interested in identifying “elite” users. and NGOs—all remain well represented among highly followed users. Consequently it provides an interesting context in which to address Lasswell’s maxim. [3] compared three measures of influence—number of followers. the kinds of effects that are of most interest to communications theorists. radio. the top ten most-followed users on Twitter are not corporations or media organizations. and understanding their role in introducing information into Twitter. mostly celebrities. these features are in general poor predictors of future cascade size. outline two different sampling methods. particularly who pays attention to whom. Kwak et al. who we differentiate from “ordinary” users in terms of their visibility. Cha et al. And finally.while. but individual people. and number of mentions— and also found that the most followed users did not necessarily score highest on the other measures. represents the full spectrum of communications from personal and private to “masspersonal” to traditional mass media. in the opposite direction interpersonal communication has become increasingly amplified through personal blogs. Moreover.ac. governments. [15] compared number of followers and page rank with a modified page-rank measure which accounted for topic. Finally. [8]. in spite of these shifts away from centralized media power. authors. To this end. and are often extremely active. organizations. that is. Kwak et al. number of retweets. including section 3.5B edges. journalists. we used the follower graph studied by Kwak et al. remain difficult to measure on Twitter.

news organization vs. Lady Gaga appears on nearly 140. as practiced by ordinary individuals communicating with their friends. And second. 2009. we note that almost all URLs broadcast on Twitter have been shortened using one of a number of URL shorteners. but the maximum # friends is several hundred thousand. brandonacox. Under the modest assumption of 40M users (roughly the number in the 2009 crawl by [8]). because URLs point to online content outside of Twitter. Finally. for example. 3. mashable. Gizmodo. etc.1 Snowball sample of Twitter Lists The first method for identifying elite users employed snowball sampling. the out-degree distribution is even more skewed than the in-degree distribution. Rather than pursuing a strategy of automatic classification.2 Twitter Firehose In addition to the follower graph.g. we are also interested in the relationships between these categories of users. our proposed classification method must also satisfy a practical constraint—namely that the rate limits established by Twitter’s API effectively preclude crawling all lists for all Twitter users3 . jsinkeywest. 2010 using data from the Twitter “firehose. Whole Foods • Blogs4 : BoingBoing. but instead resembles more the mixture of one-way mass communications and reciprocated interpersonal communications described above. does not conform to the usual characteristics of social networks.). and (b) their lack of overlap between categories: The Twitter API allows only 20K calls per hour. the name of the list can be considered a meaningful label for the listed users. In addition. and decide whether the new list is public (anyone can view and subscribe to this list) or private (only the list creator can view or subscribe to this list). the follower graph is also characterized by extremely low reciprocity (roughly 20%)—in particular. in other words. First.000) 4 The blogger category required many more seeds because bloggers are in general lower profile than the seeds for the other categories 3 3. In addition. bbrian017. hishaman. List creation therefore effectively exploits the “wisdom of crowds” [12] to the task of classifying users. where at most 20 lists can be retrieved for each API call. In particular. NotAProBlog. kbloemendaal. we restrict our attention to four classes of what we call “elite” users: media. extremejohn. but it also likely underestimates the real time quite significantly. organizations.g. BlazingMinds. where both approaches have advantages and disadvantages.3 Twitter Lists Our method for classifying users exploits a relatively recent feature of Twitter: Twitter Lists. TycoonBlogger. In additional to these theoretically-imposed constraints. discussed below. Ileane. danielscocco. shoemoney. FamousBloggers. Twitter notation for how many others a user follows). therefore. allowing us to observe when a particular piece of content is either retweeted or subsequently reintroduced by another user. For each of the four categories above. the median is less than 100.” the complete stream of all tweets2 . Chrisbrogan. as well as the relationships between these elite users and the much larger population of “ordinary” users. the most-followed individuals typically do not follow many others. we are interested in the content being shared on Twitter—particularly URLs—and so we examined the corpus of all 5B tweets generated over a 223 day period from July 28. From the total of 5B tweets recorded during our observation period. where each user is included on at most 20 lists. Paris Hilton • Media: CNN. however. bloggersblog. are interested in the relative importance of mass communications. smartbloggerz. of which the most popular is http://bit. Element321. URLs add easily identifiable tags to individual tweets. we 2 http://dev. virtuosoblogger. Before describing our methods for classifying users in terms of the lists on which they appear. we chose a number u0 of seed users that were highly representative of the desired category and appeared on many category-related lists.ly/. celebrities. engadget. for two reasons. 3. Yahoo! Inc. therefore. In both friend and follower distributions. dragonblogger. In particular. remarkablogger. the user can add/edit/delete list members. celebrity. Clearly this time could be reduced by deploying multiple accounts. Since its launch on November 2. Thus we instead devised two different sampling schemes—a snowball sample and an activity sample— each with some advantages and disadvantages. and interpersonal communications. it is useful for us to restrict attention to tweets containing URLs. this would require 4 ∗ 106 /2 ∗ 103 = 2.. 000 hours. while a small number of users have millions of followers. Because our objective is to understand the flow of information.twitter.ly URLs. and also how they are perceived (e. problogger. Lady Gaga. 2009 to March 8. To create a Twitter List. wchingya. The Twitter follower graph. motivated by theoretical arguments such as the theory of the two-step flow [6]. ditesco After reviewing the lists associated with these seeds. a user needs to provide a name (required) and description (optional) for the list. which exhibit much higher reciprocity and far less skewed degree distributions [7]. we emphasize that we are motivated by a particular set of substantive questions arising out of communications theory. As the purpose of Twitter Lists is to help users organize users they follow. we focus our attention on the subset of 260M containing bit. copyblogger. as practiced by media and other formal organizations. World Wildlife Foundation. the following keywords were hand-selected based on (a) their representativeness of the desired categories. For each category. they provide a much richer source of variation than is possible in the typical 140 character tweet. kikolani.com/doc/get/statuses/firehose .of in-degree (# followers) and out-degree (# “friends”. both in terms of their importance to the community (number of lists on which they appear). JimiJones. Once a list is created. the following seeds were chosen: • Celebrities: Barack Obama. GrowMap. Twitter Lists have been welcomed by the community as a way to group people and organize one’s incoming stream of tweets by specific sets of users. or 11 weeks.3. and bloggers. as many users appear on many more than 20 lists (e. seosmarty. masspersonal communications as practiced by celebrities and prominent bloggers. New York Times • Organizations: Amnesty International. our approach depends on defining and identifying certain predetermined classes of theoretical interest.

we not only need to classify users into categories. the bias is likely to be quite different from any introduced by the snowball sample.685 users. 524. “celebs”. In total.celebrities-ontwitter. thus obtaining similar results from the two samples should give us confidence that our findings are not artifacts of the sampling procedure. but only the latter two lists would be kept after pruning.614 of the activity sample. 000 lists.8% 216.3. but also arrive at a definition of “elite” that satisfies a tradeoff between (a) keeping each category relatively small. 116 users were obtained. bloggers Having selected the seeds and the keywords for each category.7% 127.3. charities. Importantly. hollywood. 97. In particular. however. This “activity-based” sampling method is also clearly biased towards users who are consistently active. as evidenced by the relatively flat slope of the attention curves. Specifically. 3.3% 14.2 Activity Sample of Twitter Lists Although the snowball sampling method is convenient and is easily interpretable with respect to our theoretical motivation. where all remaining users are hence- . therefore.186 35.3% 524. we first rank all users in each of category by how frequently they are listed in that category. however. Next. causes. we also generate a sample of users based on their activity. and that bloggers are more heavily represented among the activity sample at the expense of the other three categories— consistent with our claim that the two methods introduce different biases.” To resolve this ambiguity. The number of users assigned in this manner to each category is reported in Table 1.116 100% Activity Sample # of users % of users 14. we obtained a much-reduced sample of 113.e. For instance. The resulting “list of lists” was then pruned to contain only the l0 lists whose names matched at least one of the chosen keywords for that category. we crawled all lists on which that seed appeared. Also in both cases. To this end. brands. media. organizations. cause. Balancing the requirements described above. for both sampling methods. stars. To address this concern. In order to identify categories of elite users. we measure the flow of information from the top k users in each of the four categories to a random sample of 100K ordinary (i. we then performed a snowball sample of the bipartite graph of users and lists (see Figure 1). We then assigned each user to the category in which he or she has the highest membership score. Figure 2(a) shows for the snowball sample the share of following links (square symbols) and tweets received (diamonds) by an average user. celebrity. and then repeated these last two steps to complete the crawl. it is also potentially biased by our particular choice of seeds. however. the bulk of the attention is accounted for by a relatively small number of users within each category.778 13. 000.u0 l0 u1 l1 u2 l2 Table 1: Distribution of users over categories category celeb media org blog total Snowball Sample # of users % of users 82. celebrities outrank all other categories. who appeared on 7. celebrities. organisations. In addition. We note that the number of lists obtained by the activity sampling methods is considerably smaller than that obtained by the snowball sample. the two sets of results are qualitatively similar. Lady Gaga is on lists called “faves”. many of the more prominent users appeared on lists in more than one category—for example Oprah Winfrey is frequently included in lists of “celebrity” as well as “media. Interestingly. we crawl all lists associated with all users who tweet at least once every week for our entire observation period. companies. where we note that the curve for celebrities asymptotes more slowly than for the other three categories. blogs.853 18. This method initially yielded 750k users and 5M lists. organizations. news-media • Organizations: company.685 100% Figure 1: Method Schematic of the Snowball Sampling • Celebrities: star. where Table 1 reports the number of users assigned to each category.483 24. products. and assigning users to unique categories (as described above).891 13. organisation. celebs. corporation.010 41. and “celebrity”.770 15.1% 43. and the proportion of tweets the user received from everyone the user follows in each category. while Figure 2(b) shows the same information for the activity sample. or 85%. so as to facilitate comparisons. followed by the media.830 38. ngo • Blogs: blog. we computed a user i’s membership score in category c: nic . celebrity-tweets • Media: news. we chose k = 5000 as a cut-off for the elite categories. finding all users that appeared in the “celebrity” list with Lady Gaga).3 Classifying Elite Users 3. so as not to include users who are not distinguishable from ordinary users. blogger. it is also desirable to make the four categories the same size. organization. wic = Nc where nic is the number of lists in category c that contain user i and Nc is the total number of lists in category c. while (b) maximizing the volume of attention that is accounted for by each category.2% 97. For each seed. after pruning the lists to those that contained at least of the keywords above. and bloggers.6% 113. charity. Although the numerical values differ slightly. also appear in the snowball sample. We then crawled all u1 users appearing in the pruned “list of lists” (for instance. however. unclassified) users in two ways: the proportion of people the user follows in each category.0% 40. suggesting that the two sampling methods identifying similar populations of elite users–as indeed we confirm in the next section. celebsverified. celebrity-list.

“organization”. Equally interesting.228. The one slight exception to this rule is that organizations pay more attention to bloggers than to them- . and ollehkt is the Twitter account for KT. noting that both methods generate similar results. Table 2 shows that although ordinary users collectively introduce by far the highest number of URLs. “blog”. The prominence of elite users also raises the question of how these different categories listen to each other. Mashable and ProBlogger are both prominent US blogging sites. users classified as “media” easily outproduce all other categories. organizations. and Asahi. we compute the volume of tweets exchanged between elite categories. Table 3.10 friends tweets received 0 0 1000 4000 7000 top k organizations 10000 1000 4000 7000 top k blogs 10000 average % 10 20 30 0 1000 4000 7000 top k 10000 0 average % 10 20 30 friends tweets received friends tweets received 1000 4000 7000 top k 10000 (a) Snowball sample celebrities average % 10 20 30 average % 10 20 30 media friends tweets received dle for actor Ashton Kusher. only about 15% of tweets received by ordinary users are received directly from the media. for example.74 blog 1. In addition. “media”. ordinary users on Twitter are receiving their information from many thousands of distinct sources.81 media 5.05% of the user population. In the rest of this paper. Google. “media”. while Kibe Loco and Nao Salvo are popular blogs in Brazil. In the media category. celebrities average % 10 20 30 average % 10 20 30 media friends tweets received Table 2: # of URLs initiated by category # of URLs category # of URLs per-capita celeb 139. it remains the case that 20K elite users. among the blogging category. and Twitter are obviously large and socially prominent corporations.739 1023. compared with over 1. “WHO LISTENS TO WHOM” (b) Activity sample Figure 2: Average fraction of # following (blue line) and # tweets (red line) for a random user that are accounted for by the top K elites users crawled Based on this definition of elite users. we refer the top 5K users drawn from the snowball sample listed as “celebrity”. CNN Breaking News and the New York Times are most prominent.131 272. Finally. is that in spite of this fragmentation. attracts almost 50% of all attention within Twitter. information flows have not become egalitarian by any means.forth classified as ordinary. Among the celebrity list. exhibiting striking homophily with respect to attention: celebrities overwhelmingly pay attention to other celebrities. one of the first celebrities to embrace Twitter and still one of the most followed users. Figure 3 shows the average percentage of tweets that category i receives from category j.000 users identified by the sampling method. however. followed by Breaking News. Ellen Degeneres. while the remain celebrity users—Lady Gaga.5M followers. Specifically. a leading Japanese daily newspaper. Time. are all household names. To address this issue. media actors pay attention to other media actors. Table 3: Top 5 users in each category Celebrity Media Org Blog aplusk cnnbrk google mashable ladygaga nytimes Starbucks problogger TheEllenShow asahi twitter kibeloco taylorswift13 BreakingNews joinred naosalvo Oprah TIME ollehkt dooce friends tweets received 0 1000 4000 7000 top k organizations 10000 0 1000 4000 7000 top k blogs 10000 average % 10 20 30 0 1000 4000 7000 top k 10000 0 average % 10 20 30 friends tweets received friends tweets received 1000 4000 7000 top k 10000 4. comprising less than 0. Ordinary users originate on average only about 6 URLs each. followed by bloggers. therefore. from this point on. Among organizations. formerly Korean Telecom. we restrict our analysis to elite categories to the top 5. and dooce is the blog of Heather Armstrong. “blog”. “aplusk. In particular.360. suggests that the sampling method yields results that are consistent with our objective of identifying users who are prominent exemplars of our target categories. which shows the top 5 users in each of the four categories. a widely read “mommy blogger” with over 1. and celebrities. members of the elite categories are far more active on a per-capita basis. while JoinRed is the charity organization started by Bono of U2. respectively. “organization”.698 104. Starbucks. Oprah Winfrey. Clearly.058 27.364 6.000 for media users. when we talk about “celebrity”. and Taylor Swift. most of which are not traditional media organizations—even though media outlets are by far the most active users on Twitter.” is the han- The results of the previous section provide qualified support for the conventional wisdom that audiences have become increasingly fragmented.94 org 523.119.03 ordinary 244. Even if the media has lost attention relative to other elites. and so on.

Nevertheless. the population comprises two types—those who receive essentially all of their media-originating information via two-step flows and those who receive virtually all of it directly from the media. what proportion is broadcast directly to the masses. is that even users who received up to 100 media URLs durAs before. or making use of an informal convention such as “RT @user” or “via @user. in what respects they differ from other ordinary users. What is surprising. where of the 1M total. however. In general. the former type is exposed to less total media than the latter. where a user subsequently tweets a URL that has previously been introduced by another user.46 therefore represents the proportion of mediaoriginated content that reaches the masses via an intermediary rather than directly. as supposed by early theories of mass communication. For each member of this 600K subset we then counted the number n2 of these URLs that they received via non-media friends. along with an explicit acknowledgement of the source—either using official retweet function provided by Twitter. performing this analysis for the entire population of over 40M ordinary users proved to be computationally unfeasible. and for each user. 5 4. To quantify the extent to which ordinary users get their information indirectly versus directly from the media. As with our previous measure of attention. Unsurprisingly. and if the latter.” The second mechanism is what we label reintroduction. however. Figure 3. and what proportion is transmitted indirectly via some population of intermediaries? In addition. via a two-step flow. therefore. but without the acknowledgment.1 retweets per person for ordinary users—the total number of URLs retweeted by bloggers (465k) is vastly outweighed by the number retweeted by ordinary users (46M). or reintroducing it. a particular weak measure of attention for the simple reason that many tweets go unread. as claimed by the two-step flow theory. even though on a per-capita basis bloggers disproportionately occupy the role of information recyclers—93 retweets per person. but a user received the URL from another source. However. having discovered the URL outside of Twitter. A stronger measure of attention. arguably the theory that has most successfully captured the dueling importance of mass media and interpersonal influence. is to consider instead only those URLs introduced by category i that are subsequently retweeted by category j. light on the theory of the two-step flow. it should be noted. in which case we assume the information has been rediscovered independently. The first is via retweeting.1 Two-Step Flow of Information Examining information flow on Twitter can also shed new . to the extent they exist. this average is somewhat misleading. are drawn from other elite categories or from ordinary users. compared to only 1. we treat retweets and reintroductions equivalently.Celeb Category of Twitter Users A B Media B receive tweets from A Org Blog Figure 3: Share of tweets received among elite categories Figure 4: RT behavior among elite categories selves. we may inquire whether these intermediaries. on Twitter the flow of information to the masses from the media accounts for only a fraction of the total volume of information. If the first occurrence of a URL in Twitter came from a media user. As Figure 5 shows. retweeting is strongly homophilous among elite categories. As we have already noted. For the purposes of studying when a user receives information directly from the media or indirectly through an intermediary. The average fraction n2 /n = 0. shows only how many URLs are received by category i from category j.ly URLs they had received that had originated from one of our 5K media users. and which to ignore. we note that there are two ways information can pass through an intermediary in Twitter. that is. counted the number n of bit. however. it is still a substantial fraction. whether they are citing the source within Twitter by retweeting the URL. In reality. we sampled 1M random ordinary users5 . in fact. This result reflects the characterization of bloggers as recyclers and filters of information. thus their overall impact is relatively minimal. 600K had received at least one such URL. attention paid by organizations is more evenly distributed across categories than for any other category. Before proceeding with this analysis. which occurs when a users explicitly rebroadcasts a URL that he or she has received. Figure 4 shows how much information originating from each category is retweeted by other categories. so it is still interesting to ask: for the special case of information originating from media sources. but passes first through an intermediate layer of “opinion leaders” who decide which information to rebroadcast to their followers. The essence of the two-step flow is that information passes from the media to the masses not directly. bloggers are disproportionately responsible for retweeting URLs originated by all categories. then that source can be considered an intermediary.

Figure 6: Frequency of intermediaries binned by # randomly sampled users to whom they transmit media content. Bakshy et al. U. World.g. only 9 categories had more than 100 URLs during the observation period. by definition. [6]. NY Times. [1]. like their followers. sports news. is roughly ten times as active as CNN Breaking News. and on every social and economic level. Given the length of time that has elapsed since the theory of the two-step flow was articulated. Of the 6398 New York Times bit. Of these. however.. however.” corresponding to our classification of most intermediaries as ordinary. after CNN Breaking News. etc.000 URLs along a variety of dimensions. Arts. Figure 7 shows the proportion of URLs from each New York Times section retweeted or reintroduced by each category. etc. thus we focused our attention on the remaining 8 topical categories. for example acts as an intermediary for over 100. the population of intermediaries is smaller than that of the users who rely on them. but also show clear differences in how the attention is allocated to the different elite categories. however. The original theory also emphasized that opinion leaders. just like other ordinary users. not elites. that a service like Twitter was likely unimaginable at the time—it is remarkable how well the theory agrees with our observations. Sports.nytimes. ing our observation period received all of them from opinion leaders.a b 10000 1 0 100 c d # of opinion leaders 4 16 64 256 2048 16384 131072 # of two−step recipients Figure 5: Percentage of information that is received via an intermediary as a function of total volume of media content to which a user is exposed. one of which—“NY region”—was highly specific to the New York metropolitan area. Finally. WHO LISTENS TO WHAT? The results in section 4 demonstrate the “elite” users account for a substantial portion of all of the attention on Twitter. versus 7).000 users.html?ref=category . is the second-most-followed news organization on Twitter. communications technology in the interim—given.5M followers. the vast majority of which (96%) are classified ordinary users.) 6 . and the transformational changes that have taken place in 5. hence the number of intermediaries who rely on two-step flows is much smaller than for random users. entertainment. World 6 http://www. where we note that the most prominent intermediaries are disproportionately drawn from the 4% elite users—Ashton Kucher (asplusk). for example.S. we restrict attention to URLs originated by the New York Times which. Finally. a minority satisfies this function for multiple users. but rather a continuously varying one. Interestingly. and the many different ways one can classify content (video vs. Figure 6 shows that although all intermediaries.). these results are all broadly consistent with the original conception of the two-step flow. in fact. Science. also received at least some of their information via two-step flows. Given the large number of URLs in our observation period (260M ). news vs. Comparing Figure 5a and 5c. corresponding to our finding that intermediaries vary widely in the number of users for whom they act as filters and transmitters of media content. advanced over 50 years ago. we note that intermediaries are not like other ordinary users in that they are exposed to considerably more media than randomly selected users. but still surprisingly large.com/year/month/day/category/ title. with over 2. we exploit a convenient feature of their format—namely that all NY Times URLs are classified in a consistent way by the section in which they appear (e. and how many of them are there? In total. however. political news vs. used Amazon’s Mechanical Turk to classify a stratified sample of 1. Who are these intermediaries. classifying even a small fraction of URLs according to content is an onerous task. 6370 could be successfully unshortened and assigned to one of 21 categories. text. this method does not scale well to larger sample sizes. but that in general they were more exposed to the media than their followers—just as we find here. In addition. which emphasized that opinion leaders were “distributed in all occupational groups. pass along media content to at least one other user.ly URLs we observed. Instead. Figure 5c also shows that at least some intermediaries also receive the bulk of their media content indirectly. roughly 500K. To classify NY Times content. so it is arguable a better source of data. Interestingly. It is therefore interesting to consider what kinds of content is being shared by these categories. we find that on average intermediaries have more followers than randomly sampled users (543 followers versus 34) and are also more active (180 tweets on average. the theory predicted that opinion leadership was not a binary attribute.

thus for the following analysis. Technology Figure 8: Schematic of lifespan estimation procedure create large biases for long lived URLs. the more likely it will re-appear in the future. and only count URLs as having a lifespan of τ if (a) they do not appear in the first δ days. and possibly world news.15 0. Health 0. (b) they first appear in the interval between the buffers. we now explore the lifespan of URLs introduced by different categories of users. the sorts of URLs that are picked up by bloggers are of more persistent interest. What we observe as the lifespan of a URL.1. Arts 6. however. we need to pick τ and δ such that τ + 2δ ≤ 223. To determine δ we first split the 223 day period into two segments—the first 133 day estimation period and the last 90 day evaluation period (see Figure 8(b))—and then ask: if we (a) observe a URL first appear in the first 133−δ days and (b) do not see it in the δ days prior to the splitting point. In other words. measuring lifespan seems a trivial matter.10 0. we seek to determine a buffer δ at both the beginning and the end of our 223day period. as shown in Figure 9(a). and disproportionately high interest in science. as illustrated in Figure 8(a).35 0.15 0. and second. In particular. URLs initiated by the elite categories exhibit a similar distribution over lifespan to those initiated by ordinary users. therefore. Although this limitation does not create much of a problem for short-lived URLs—which account for the vast majority of our observations—it does Having established a method for estimating URL lifespan. URLs introduced by different types of elite users or ordinary users may exhibit different lifespans. technology. where increasingly niche categories like Health. Using this estimation/evaluation split.00 5. how likely are we see it in the last 90 days? Clearly this depends on the actual lifespan of the URL. as the longer a URL lives. the overall pattern is replicated for all categories of users. User Category Figure 7: Number of RTs and Reintroductions of New York Times stories by content category news is the most popular category. when looking at the percentage of URLs of different lifespans initiated by each category.20 2. In general. we consider only URLs that have a lifespan τ ≤ 70. and Sports. those that only appeared once). while the media shows somewhat greater interest in U.00 blog celeb media org other blog celeb media org other 8. Arts.05 0.00 3. Both of these results can be explained by the type of content that originates from different sources: whereas news stories tend to be replaced by updates on a daily or more frequent basis. Sports δ Total observation window = 223 days % RTs and Re-introductions 0.05 0.10 0.10 0. and because we can only classify a URL as having lifespan τ if it appears at least τ days before the end of our window. News.15 0.25 0. a URL that is last observed towards the end of the observation period may be retweeted or reintroduced after the period ends.20 0. Naively. we see two additional results: first. show greater interest in sports and less interest in health. Science.30 0. News first observation of URL last observation of URL δ τ δ estimation period = 133 days evaluation period = 90 days τ 4. because we require a beginning and ending buffer. U. and so are more likely to be retweeted or reintroduced months or even years after their initial introduction. and Technology are less popular still. World News 0. URLs that appear towards the end of our observation period will be systematically classified as shorter-lived than URLs that appear towards the beginning.10 0.2 Lifespan By Category 5. by contrast.S.00 7. 5.30 0. organizations show disproportionately little interest in business and arts-related stories.30 0. To address the censoring problem.15 0.S. URLs originated by bloggers are overrepresented among the longer-lived content.25 0.S.35 0. Finally.05 0.25 0.25 0. We determined that τ = 70 and δ = 70 sufficiently satisfied our constraints. a finite observation period—which results in censoring of our data—complicates this task. news stories.35 0. we find an upper-bound on lifespan for which we can determine the actual lifespan with 95% accuracy as a function of δ.30 0. however.35 0. Business.20 0. is in reality a lower bound on the lifespan. by which we mean the time lag between the first and last appearance of a given URL on Twitter. a URL that is first observed toward the beginning of the observation window may in fact have been introduced before the window began. followed by U.05 0. As Figure 9(b) shows. Business 0. while correspondingly. Science 0. URLs originated by media actors generate a large portion of short-lived URLs (especially URLs with lifespan=0.1 Lifespan of Content In addition to different types of content.20 0. . and (c) they do not appear in the last δ days. but there are some minor deviations: in particular. Celebrities.

movie clips. For the vast majority of URLs on Twitter. for ordinary users.6 0. and longformat magazine articles have lifespans that are effectively unbounded. grouped by the categories that introduced the URL7 . music. for URLs introduced by elite users.0 0. after which a given story is unlikely to be reintroduced or rebroadcast. even for URLs that persist for weeks. consistent with our interpretation above. Two related points are illustrated by Figure 11.20 log(# of URLs with lifespan = x day) q q q 15 q q q qq q q other celeb media org blog q qq qq q qq 10 qq qqqqqqq qqqq qqq qqqqq qqqqq qqq qqqqq qqqqqqq qqq qqqq qqq qqq 5 0 0 10 20 30 40 lifespan (day) 50 60 70 Figure 10: Top 20 domains for URLs that lived more than 200 days RT rate (# of RTs / total # of occurrences) (a) Count 1. CONCLUSIONS In this paper. the source of its origin also impacts its persistence. we used the bit. thus the RT rate is zero.0 q q qq q qq qqqq q q q qqqqqqqq qq qqqqqqqqqqqqqqq qqq qqq qqq qqq qqq qqqqqq qqq q q q q qq q 0 10 20 30 40 50 60 70 70 lifespan (day) (b) Percent Figure 9: 9(a) Count and 9(b) percentage of URLs initiated by 5 categories. and mapped them into 21034 web domains. captured by the first part of Laswell’s maxim—“who says what to whom”—in the context that only appeared once in our dataset. in other words. the result is somewhat the opposite—that is. however. longevity is determined not by diffusion. Second. with different lifespans Figure 11: Average RT rate by lifespan for each of the originating categories of appearances of URLs after the initial introduction derives not from retweeting.2 0. Note here that URLs with lifespan = 0 are those URLs . the population of long-lived URLs is dominated by videos. First. Twitter is. Some of this content—such as daily news stories— has a relatively short period of relevance. should be viewed as a subset of a much larger media ecosystem in which content exists and is repeatedly rediscovered by Twitter users. and books.ly API service to unshorten 35K of the most long-lived URLs (URLs that lived at least 200 days).4 q q 7 % of URLs from elites category 6 5 4 3 2 1 0 0 10 20 30 40 lifespan (day) 50 60 celeb media org blog other celeb media org blog 0. Although it is unsurprising that elite users generate more retweets than ordinary users. To shed more light on the nature of long-lived content on Twitter. which shows the average RT rate (the proportion of tweets containing the URL that are retweets of another tweet) of URLs with different lifespans. in other words. As Figure 10 shows. and suggests that in spite of the dominant result above that content lifespan is determined to a large extent by type.8 0. classic music videos. but rather from reintroduction. and can seemingly be rediscovered by Twitter users indefinitely without losing relevance. where this result is especially pronounced for long-lived URLs. At the other extreme. the size of the difference is nevertheless striking. we investigated a classic problem in media communications research. the majority 7 6. but by many different users independently rediscovering the same content. at least on average—a result that is consistent with previous findings [1]. they are more likely to be retweeted than reintroduced.

. Social theory and social structure. shedding more light on the “what” element of Lasswell’s maxim. Finally. 2006. possibly identifying opinion leaders for different topics. Bakshy. 1968. 3rd edition. State University of New York Press. 1968. for example. we find considerable support for the two-step flow of information—almost half the information that originates from the media passes to the masses indirectly via a diffuse intermediate layer of opinion leaders. F. Science. attention remains highly concentrated. 1st edition. What is twitter. Kossinets and D. devote a surprisingly small fraction of their attention to business-related news. a significant challenge for future work is to merge the data regarding information flow on Twitter with other sources of outcome data—relating. we find that the longest-lived URLs are dominated by content such as videos and music. Choi. The people’s choice. Media sociology: The dominant paradigm. Gaudet. 58(4):707–731. economies. and bloggers following bloggers. and H. S. and D. J. Measuring user influence on twitter: The million follower fallacy. D. and Q. Routledge. such as TV and radio on the one hand. A new era of minimal effects? the changing foundations of political communication. pages 17–38. P. Albany. [10] P. 1948. Lazarsfeld. S. [5] T. editor. ” [7] G. J. Second. Journal of Communication. and Culture on Social Network Sites. celebrities. C. 2010. Gitlin. and B. etc). 1978. our analysis sheds new light on some old questions of communications research. communications on Twitter may be unrepresentative of information flow via more traditional channels. S. In addition. B. Park.. Second. and interpersonal interactions on the other hand. who although classified as ordinary users. Columbia University Press.05% of the population accounts for almost half of all attention. J. Kwak. H. Personal influence. REFERENCES [1] E. Theory and Society. Weimann. Hofman. The Communication of Ideas. Glencoe. we note that although our use of Twitter lists to label users was motivated by a specific set of questions regarding mass vs interpersonal communications. Bryson. pages 261–270. L. [14] G. [4] J. 2003070095 James Surowiecki. Washington. [6] E. E. In R. In particular. Kim. are more connected and more exposed to the media than their followers. Identifying ‘influencers’ on twitter. and media influence sources online. [11] R. Benevenuto. A. NY. U. there are some differences— organizations. P. 1957. bloggers. New York. Sociometry. In closing. IL. and S. Within the population of elite users. K. Tong. University of Illinois Press. Moon. Bennett and S. [9] H. [2] W. DeAndrea. By studying the flow of information among the five categories that we identified (media. Surowiecki. Merton. ACM. D. media-originated URLs are disproportionately represented among short-lived URLs while those originated by bloggers tend to be overrepresented among long-lived URLs. with celebrities following celebrities. [15] J. A Networked Self: Identity. Community. pages 441–474. In L. Walther. 7. Van Der Heide. By restricting our attention to Twitter. the part played by people in the flow of mass communications. C. Urbana.S. it provides an unprecedented level of resolution and coverage regarding who is listening to whom. Gummad. Business. Lasswell. for example. peer. Empirical analysis of an evolving social network. Free Press. where roughly 0. we feel the advantages of using Twitter to answer this question outweighed the limitations. 2010. And finally. and H. which are continually being rediscovered by Twitter users and appear to persist indefinitely. societies. Lee. and K. ACM. T. Katz and P. Hong Kong. W. Lim. Papacharissi. our conclusions are necessarily limited to one narrow cross-section of the media landscape. Free Press. ACM. Doubleday. editor. B. another area for future work would be to extract content information in a more systematic manner. Moreover. crowdsourced classification scheme of users. Watts. 1955. The Influentials: People Who Influence People. [3] M. editor. Patterns of influence: Local and cosmopolitan influentials. Jiang. M. 6(2):205–253. In Fourth ACM International Conference on Web Seach and Data Mining (WSDM). In particular. The diffusion of an innovation among physicians. However. 2011. organizations. to opinions or actions that would engage more directly with the “effects” component of Lasswell’s maxim. it would also be interesting to explore automatic classification schemes from which additional user categories could emerge. First. Twitter effectively provides a ready-made. S. We also find that different types of content exhibit very different lifespans. as has been proposed elsewhere [14]. New York. pages 591–600. Ill. we find that although all categories devote a roughly similar fraction of their attention to different categories of news (World. Twitterrank: finding topic-sensitive influential twitterers. Cha.of Twitter. 20(4):253–270. and because Twitter maintains a complete record of every tweet broadcast. In Z. Interaction of interpersonal. moreover. Haddadi. Berelson. E. 2004. Winter. and ordinary). DC. a social network or a news media? In Proceedings of the 19th international conference on World Wide Web. media following media. C. Includes bibliographical references. The Wisdom of Crowds : Why the many are smarter than the few and how collective wisdom shapes business. J. 311(5757):88–90. [12] J. F. The structure and function of communication in society. and that for this reason we have focused on a limited set of predetermined user-categories. In 4th Int’l AAAI Conference on Weblogs and Social Media. Third. F. 1994. 2010. Coleman. such an approach would allow one to examine the category of opinion leaders in more detail. 2008. pages 117–130. how the voter makes up his mind in a presidential campaign. attention is highly homophilous. Iyengar. Lazarsfeld. Merton. Katz. and nations. In Proceedings of the third ACM international conference on Web search and data mining. because Twitter users themselves classify other users by including them on lists. Menzel. Mason. J. New York. Weng. T. First. K. [8] H. because Twitter users explicitly opt-in to “follow” each other. Watts. H. Carr. He. [13] J. we find that although audience attention has indeed fragmented among a wider pool of content producers than classical models of mass media. 2010.