Social Listening-Driven Product Innovation in the Coffee Market: A Twitter-Based Framework for Consumer Preference Modeling and Importance–Satisfaction Opportunity Mapping
Abstract:
Rapidly changing consumer preferences, consumption occasions, service experiences, and promotional practices require coffee businesses to continuously identify and evaluate emerging innovation opportunities. However, converting large volumes of unstructured social media discourse into actionable knowledge remains a major challenge, particularly for businesses with limited market research capabilities. This study investigates how social listening can support technology-enabled market sensing and product innovation in the coffee industry. Four Twitter/X corpora collected during 2023 were analyzed, comprising a global coffee corpus of 172,017 tweets and three brand-specific corpora relating to Kopi Kenangan, Starbucks, and Kopi Janji Jiwa. A Biterm Topic Model (BTM) was applied to identify latent consumer preference structures, while topic prevalence and sentiment-derived satisfaction were integrated through importance–satisfaction mapping to distinguish market-supported innovation opportunities. The global corpus yielded 13 preference topics, five of which represented opportunities associated with work-related benefits, coffee-shop experiences, drinking enjoyment, caffeine and sleep, and morning consumption. The analysis identified three opportunity topics among eight preference topics for Kopi Kenangan, five among 12 topics for Starbucks, and ten among 15 topics for Kopi Janji Jiwa. Cross-brand comparison showed that innovation opportunities extended beyond beverage attributes to include service experience, complementary products, pricing, promotions, social influence, customer segments, and collaboration channels. The findings demonstrate that social listening can function as a decision-oriented innovation mechanism rather than merely a descriptive monitoring tool. The proposed framework advances data-driven innovation research by linking digital consumer discourse to structured opportunity identification and provides a practical approach to market sensing, innovation prioritization, and low-cost experimentation for resource-constrained coffee businesses.1. Introduction
Innovation is essential to the competitiveness and long-term sustainability of the food and beverage industry, where consumer preferences change rapidly, product life cycles continue to shorten, and meaningful differentiation is increasingly difficult to sustain. These conditions are particularly evident in the coffee market because consumer value is determined by more than the intrinsic properties of the beverage. Coffee consumption encompasses multiple dimensions, including taste, milk and topping variants, price, consumption occasion, store atmosphere, service delivery, complementary foods, lifestyle associations, and social interaction (Carr et al., 2015; Samoggia et al., 2020; Vidal et al., 2015). Coffee innovation should therefore be approached as a multidimensional product–service–market challenge rather than as a matter of beverage formulation alone. As customer requirements become more diverse, coffee businesses must determine which combinations of product attributes, service elements, consumption contexts, and market offerings should be prioritized in product development. Surveys, interviews, focus groups, and consumer tests remain valuable for eliciting customer requirements, but they are generally conducted on a limited scale and require considerable time, effort, and financial resources. Online user-generated information has consequently been recognized as a complementary source of knowledge for product development (Qi et al., 2016; Zhang et al., 2022). Firms operating in rapidly changing consumer markets therefore require scalable methods that can capture emerging customer needs continuously and convert them into evidence for innovation decisions.
Digitalization creates new possibilities for innovation management because consumers increasingly use social media to communicate their opinions, experiences, preferences, complaints, and consumption routines. Social media activity encompasses interconnected functions related to identity, conversation, sharing, presence, relationships, reputation, and group formation, generating extensive and continuously evolving flows of user-generated content (Kietzmann et al., 2011). Unlike responses elicited through conventional research instruments, these communications are produced spontaneously and often describe products and services in the language consumers use in everyday settings. Social media can therefore function as both a digital voice of the customer and an external knowledge source for innovation. From an innovation-management perspective, its role has expanded beyond communication and promotion to encompass market intelligence, environmental sensing, and decision support. Geissinger et al. (2023) showed that social media analytics is increasingly used in customer-, market-, technology-, and society-oriented innovation research. This capability is particularly relevant to small and resource-constrained coffee businesses that cannot continuously conduct large-scale surveys, consumer panels, or market experiments. Within the food and beverage domain, Carr et al. (2015) demonstrated that social media analysis can complement conventional methods for examining consumer perceptions of coffee freshness, while Vidal et al. (2015) showed that Twitter conversations reveal not only what people eat or drink but also when, where, with whom, and why consumption occurs.
The availability of large volumes of social media data does not, however, automatically produce actionable innovation knowledge. Tweets are short, unstructured, heterogeneous, and frequently expressed through informal, context-dependent, or abbreviated language. Consumers rarely formulate their preferences as explicit statements identifying a particular attribute as important. Instead, their needs and evaluations are embedded in discussions of daily routines, prices, tastes, work, social activities, health concerns, retail environments, and varying levels of satisfaction. Social media analytics must consequently address substantial challenges in data discovery, collection, preparation, noise removal, and thematic interpretation (Stieglitz et al., 2018). Simple frequency-based analysis is unlikely to reveal the latent structure of consumer preferences because frequently occurring terms may not represent coherent needs, consumption contexts, or commercially meaningful opportunities. Topic modeling offers a more systematic means of identifying recurrent semantic patterns within large text collections and organizing fragmented discourse into interpretable thematic structures (Abdelrazek et al., 2023). Clusters of semantically related terms can thereby be interpreted as topics representing product attributes, consumption situations, service experiences, consumer preferences, or market mechanisms. Nevertheless, identifying what consumers discuss is only the first stage of innovation analysis. The identified topics must subsequently be evaluated to determine whether they indicate opportunities for product improvement, new product development, service innovation, market expansion, or broader business-model development.
A growing body of research has developed methods for transforming online customer information into opportunity-oriented knowledge for product development. Jeong et al. (2019), for example, proposed a social media mining approach that combined topic modeling and sentiment analysis to identify product opportunities by considering both the importance of a topic and the level of customer satisfaction associated with it. This distinction is important because a frequently discussed customer requirement does not necessarily constitute the most valuable innovation opportunity. Opportunity potential depends both on the relevance of a requirement to customer value and on the extent to which existing products or services satisfy that requirement. Subsequent research has extended this logic through more elaborate analytical designs. Choi et al. (2020) identified time-evolving product opportunities by tracing changes in customer requirements expressed on social media. In parallel, Wang et al. (2023) integrated sentiment analysis, Latent Dirichlet Allocation (LDA), topic engagement, and topic emergence to mine customer complaints and identify product opportunities. Evidence from online review research further indicates that customer requirements extracted from digital content can guide product improvement (Qi et al., 2016). The inclusion of neutral sentiment can also improve the classification and prioritization of customer requirements (Zhang et al., 2022). Purnama et al. (2023) extended online data-driven innovation by linking customer requirements identified through social media to product-, technology-, process-, and supplier-related opportunities during the early stages of new product development. Related work in automated marketing research has shown that unstructured online reviews can reveal product attributes and relative brand positions, thereby transforming the digital voice of the customer into structured market intelligence for product and marketing decisions (Lee & Bradlow, 2011).
The feasibility of this data-driven approach is supported by empirical evidence from the coffee sector and other food-related domains. Samoggia et al. (2020) examined Twitter conversations concerning coffee and health and found that digital discourse encompasses functional, hedonic, lifestyle, wellness, and consumption-related meanings. Coffee is therefore perceived not only as a physical beverage but also as part of everyday routines, energy management, pleasurable consumption, well-being, and lifestyle expression. At the brand level, Shirdastian et al. (2019) analyzed more than two million Starbucks-related tweets to investigate brand authenticity and consumer sentiment, demonstrating how large-scale social media data can be used to identify consumer perceptions and evaluations of a brand. This perspective is consistent with customer-based brand equity theory, according to which consumer responses are shaped by brand knowledge, awareness, image, and associated meanings (Keller, 1993). Comparable evidence has also emerged outside the coffee sector. Feldmeyer & Johnson (2022) applied Twitter data and topic modeling to examine consumer perceptions and product-development opportunities relating to turmeric. Their findings indicated that spontaneous consumer discourse can reveal contexts of use, product combinations, lifestyle associations, and unexplored areas for innovation. Despite this progress, most food and beverage studies remain primarily concerned with consumer perceptions, sentiment, health-related attributes, consumption contexts, or information associated with a single brand. Fewer studies establish a direct analytical pathway from digital consumer discourse to the systematic identification and prioritization of innovation opportunities.
Several research gaps therefore remain. First, although previous research has established the value of social media for understanding consumer perceptions of food and coffee (Carr et al., 2015; Samoggia et al., 2020; Shirdastian et al., 2019; Vidal et al., 2015), limited attention has been given to the systematic transformation of coffee-related online discourse into product innovation opportunities. Existing studies reveal what consumers discuss and how they evaluate products, but they do not always translate these findings into structured priorities for innovation decisions. Second, established opportunity-mining frameworks incorporate measures such as importance, satisfaction, sentiment, complaints, and changes in customer requirements over time (Choi et al., 2020; Jeong et al., 2019; Qi et al., 2016; Wang et al., 2023; Zhang et al., 2022), yet these approaches have rarely been applied to innovation opportunities in the coffee business. Third, food-related social media studies commonly examine either a broad product category or an individual brand. Research that simultaneously compares category-level discourse with conversations surrounding multiple brands remains limited, even though these contexts may generate substantially different innovation signals. General coffee discourse may emphasize consumption routines, functional benefits, or lifestyle meanings, whereas brand-specific communities may place greater emphasis on taste variants, pricing, store experiences, promotions, complementary products, customer segments, or partnerships. Finally, consumer preference identification and product opportunity evaluation are frequently treated as separate analytical tasks. A more integrated framework is required to connect social listening, consumer preference modeling, product opportunity identification, cross-brand comparison, and managerial decision support.
This study develops a social listening-driven framework for identifying consumer preferences and transforming them into product innovation opportunities in the coffee market. The framework converts digital consumer discourse into three connected outputs: consumer preference topics, market-supported opportunity topics, and actionable directions for product, service, promotional, and channel innovation. The empirical analysis draws on four Twitter/X corpora collected during 2023: a global coffee corpus comprising 172,017 tweets and three brand-specific corpora relating to Kopi Kenangan, Starbucks, and Kopi Janji Jiwa. The cross-brand design makes it possible to examine whether category-level coffee preferences are reproduced within individual brand communities and whether particular communities generate distinct opportunity structures. Topic modeling is used to organize unstructured consumer discourse into latent preference themes, after which topic prevalence and sentiment-derived satisfaction are incorporated into importance–satisfaction mapping to identify market-supported opportunities. Rather than treating social listening as a descriptive monitoring activity, the framework positions it as a technology-enabled market-sensing and innovation-prioritization mechanism.
The study makes methodological, empirical, and managerial contributions. Methodologically, it connects social listening and consumer preference modeling with importance–satisfaction-based opportunity identification, creating an explicit analytical pathway from unstructured digital discourse to innovation priorities. Empirically, it compares market-wide coffee conversations with three brand communities, thereby revealing how opportunity structures vary across category-level and brand-specific contexts. Managerially, it translates the digital voice of the customer into decision-relevant directions concerning beverages, consumption occasions, service experiences, complementary products, pricing, promotions, customer segments, distribution channels, and collaborations. The framework is particularly relevant to resource-constrained coffee businesses because it provides a scalable approach to digital market sensing and supports the selection of innovation hypotheses that can subsequently be evaluated through limited-scale market experiments. Accordingly, the study aims to explain how social listening can be transformed into a structured mechanism for consumer preference modeling, product opportunity identification, and innovation decision support in the coffee industry.
2. Literature Review
This section establishes the theoretical and methodological foundations of social listening-driven product innovation. Section 2.1 conceptualizes social listening as a scalable source of the digital voice of the customer. Section 2.2 examines the role of social media analytics in innovation management and product decision-making. Section 2.3 reviews topic-modeling approaches for short-form consumer discourse and explains the relevance of the Biterm Topic Model (BTM). Section 2.4 considers how consumer preference topics can be converted into product opportunities through importance–satisfaction mapping. Section 2.5 reviews empirical evidence from coffee and other food-related contexts, while Section 2.6 discusses the value of digital market sensing for resource-constrained coffee businesses. Finally, Section 2.7 synthesizes the research gap and positions the contribution of the present study.
Social listening refers to the systematic observation and analysis of user-generated digital communication to identify consumer perceptions, experiences, needs, preferences, and concerns. Compared with structured surveys, social media provides access to unstructured and spontaneously generated expressions written in the language consumers use in everyday situations. Kietzmann et al. (2011) conceptualized social media through seven functional building blocks: identity, conversations, sharing, presence, relationships, reputation, and groups. These interconnected functions help explain why product-related information is distributed across numerous short communications rather than presented as complete and formally structured evaluations.
This distributed form of consumer expression creates both an analytical advantage and a methodological challenge. Its principal advantage lies in scale: large volumes of consumer discourse can be observed across products, brands, consumption situations, and periods. However, individual messages are often fragmented, informal, ambiguous, or only indirectly related to product evaluation. Computational methods are therefore required to remove noise, identify recurring patterns, and combine dispersed signals into interpretable structures. Social listening is most valuable when it complements conventional customer research by identifying emerging topics, preference patterns, and market signals that can subsequently inform design activities, controlled experiments, and market validation. It should not be treated as a substitute for direct consumer research, but as a scalable source of external knowledge that can guide the selection of questions and opportunities requiring further investigation.
Within innovation-management research, social media analytics is increasingly recognized as a means of acquiring, organizing, and interpreting external knowledge. Geissinger et al. (2023) found that social media analytics has been applied to customer-, market-, technology-, and society-oriented innovation research. Its use in innovation management nevertheless remains comparatively underdeveloped, particularly when analytical findings must be converted into specific product or market decisions. This limitation creates a need for research that moves beyond descriptive monitoring and demonstrates how social media outputs can support the identification, evaluation, and prioritization of innovation opportunities.
A data-driven approach is particularly valuable during the early stages of product development, when firms must screen a large and heterogeneous set of customer requirements, ideas, technologies, and market signals before committing resources. Purnama et al. (2023) demonstrated that online data can support opportunity discovery and accelerate the acquisition of knowledge concerning customers, technologies, processes, and supply chains. The present study extends this decision-oriented perspective to social listening in the coffee market. Online consumer conversations are treated as an external knowledge resource from which latent preference topics can be identified, evaluated, and filtered into market-supported opportunities. This progression from digital discourse to preference structures and subsequently to opportunity priorities connects social media analytics more directly with innovation decision-making.
Topic modeling comprises a family of unsupervised learning methods designed to identify latent thematic structures in large text collections. LDA, one of the foundational probabilistic topic models, represents each document as a mixture of topics and each topic as a probability distribution over words (Blei et al., 2003). Although LDA has been widely applied to textual data, social media messages present a particular methodological difficulty because their short length produces sparse within-document word co-occurrence patterns. The BTM was developed to address this limitation by modeling word-pair co-occurrences across the corpus rather than relying primarily on co-occurrences within individual documents, making it particularly suitable for short texts such as tweets (Yan et al., 2013).
The value of a topic model for innovation research depends not only on its statistical performance but also on the interpretability and decision relevance of the resulting topics. A statistically coherent topic has limited managerial value if its word distribution cannot be interpreted as a recognizable consumer need, product attribute, consumption context, service experience, or market mechanism. Topic coherence measures provide a means of comparing alternative model specifications and assessing whether the highest-probability words within a topic form a semantically meaningful group (Röder et al., 2015). The evaluation of topic models in innovation-oriented research should therefore combine quantitative evidence with domain-based interpretation. Statistical coherence indicates whether words form stable thematic structures, whereas managerial interpretation determines whether those structures can inform product, service, or market decisions.
For the empirical analysis, BTM was selected because tweets are short documents and corpus-level biterm co-occurrence reduces the sparsity that can limit document-level topic models. LDA is retained in this section as a conceptual and methodological benchmark rather than as the final analytical model. The BTM specification, candidate topic-number assessment, and final topic-selection procedure are reported in Section 3.3. This separation distinguishes the theoretical justification for selecting BTM from the operational details of its implementation.
Topic modeling reveals the subjects that consumers discuss, but innovation management requires a further assessment of which topics warrant action. Jeong et al. (2019) proposed a product opportunity-mining approach that integrated topic importance with customer satisfaction to evaluate the magnitude and direction of potential opportunities. This combined perspective provides an appropriate conceptual foundation because discussion frequency alone does not establish innovation priority. A frequently mentioned requirement may have limited strategic value if it is weakly connected to consumer value, while a strongly evaluative but rarely discussed issue may reflect only an isolated experience. Opportunity assessment must therefore consider both the relative salience of a topic and the extent to which it is associated with positive or negative consumer evaluations.
The present study uses importance–satisfaction maps for the global coffee corpus and the three brand-specific corpora. Product development opportunities are identified among topics located in the high-importance and high-satisfaction quadrant. These topics constitute a market-supported opportunity space: they are sufficiently salient within the relevant corpus and are associated with favorable consumer experiences or preferences. This interpretation differs from a conventional complaint-driven improvement matrix, in which high importance and low satisfaction indicate deficiencies requiring corrective action. In the present framework, high importance and high satisfaction indicate areas of established consumer acceptance that may support reinforcement, extension, complementary offerings, service development, communication initiatives, or channel expansion.
This distinction is particularly relevant to food and beverage innovation, where value creation is not limited to correcting dissatisfaction. Businesses can also build on attributes and experiences that consumers already value. A positively evaluated flavor profile may support a closely related product variant; a well-regarded store experience may justify service or spatial enhancements; and established interest in complementary products, promotions, or access channels may provide a basis for controlled market extension. Importance–satisfaction mapping is therefore used as an innovation-screening mechanism rather than as evidence that a commercial opportunity has already been validated. Topics identified through the map require subsequent managerial interpretation, product or service conceptualization, limited-scale experimentation, and evaluation against observable market outcomes.
Consumer discourse concerning coffee incorporates utilitarian, hedonic, experiential, and social meanings. Samoggia et al. (2020) showed that the analysis of coffee-related tweets can reveal perceptions concerning health, energy, well-being, mood, lifestyle, and consumption. Digital discussions about coffee therefore encompass not only physical product properties but also the meanings attached to particular occasions, routines, and experiences. This breadth makes social listening relevant to the identification of marketable consumption contexts and service propositions as well as sensory product attributes.
Supporting evidence is also available from other food-product domains. Feldmeyer & Johnson (2022) used Twitter data and short-text topic modeling to identify consumer perceptions and product-development white spaces relating to turmeric. Their analysis demonstrated that spontaneous consumer conversations can reveal product-use contexts, combinations with other products, lifestyle associations, and potential areas for development. Taken together, these studies indicate that social media can support food-product opportunity discovery when analytical methods move beyond word frequency and organize fragmented consumer expressions into coherent, interpretable themes.
Resource-constrained coffee businesses often lack the financial, technical, and organizational capacity to conduct recurring surveys, maintain large consumer panels, or perform repeated concept tests. In this context, social media represents a comparatively accessible source of digital market intelligence. According to Borah et al. (2022), social media use may strengthen the innovation capabilities and sustainable performance of smaller firms, and comparable findings have been reported for small businesses in Indonesia (Noviaristanti et al., 2023). Although these studies primarily examine relationships among technology adoption, innovation capability, and firm performance, they also support a broader understanding of social media as an organizational capability rather than solely as a communication or promotional channel.
The practical analytical problem confronting a resource-constrained coffee business is therefore how to reduce a large and continuously evolving body of digital conversation to a manageable set of innovation decisions. Early research on the virtual customer established that Internet-based consumer input can complement conventional methods across different stages of product development (Dahan & Hauser, 2003), while automated analysis of online reviews has shown that consumer-generated text can reveal product attributes and market structures relevant to decision-making (Lee & Bradlow, 2011). The present study extends this decision-oriented logic to the coffee market by distinguishing consumer preference topics from opportunity topics and by comparing category-level coffee discourse with conversations associated with particular brand communities. This structure allows digital market sensing to inform not only what consumers value but also which product, service, experience, promotional, or channel-related themes warrant further experimentation.
Table 1 summarizes the principal research streams relevant to the present study and clarifies its position within the existing literature. Previous research has established the value of social media analytics for innovation management, demonstrated the feasibility of topic modeling for short consumer texts, and shown that topic salience can be combined with evaluative measures to support opportunity identification. However, an operational gap remains between the collection of coffee-related social media data and its systematic transformation into consumer preference structures, market-supported opportunity topics, cross-brand comparisons, and actionable innovation directions. Existing studies typically address only part of this analytical sequence.
Study | Data/Context | Main Method | Primary Output | Relevance/Gap for this Study |
Geissinger et al. (2023) | Cross-domain innovation research | Systematic review of social media analytics | Innovation-management research agenda | Establishes the value of social media analytics for innovation but does not provide an operational opportunity-identification framework for the coffee market. |
Jeong et al. (2019) | Product-related social media | Topic modeling + sentiment/opportunity mining | Product topic importance, satisfaction, opportunity | Provides the central importance–satisfaction logic but does not examine coffee businesses or compare category-level and brand-specific communities. |
Samoggia et al. (2020) | Coffee and health on Twitter | Content and sentiment analysis | Coffee-health perceptions | Provides coffee-specific evidence but focuses on health perceptions rather than a broader portfolio of product, service, and market opportunities. |
Feldmeyer & Johnson (2022) | Turmeric tweets | Short-text topic modeling | Consumer perceptions and product white spaces | Demonstrates the use of Twitter data for food-product opportunity discovery but does not integrate cross-brand comparison or importance–satisfaction mapping. |
Borah et al. (2022) | Smaller firms | Survey/structural modeling | Social media, innovation capability, performance | Supports the relevance of social media to resource-constrained firms but does not provide an operational analytics pipeline for innovation decisions. |
This study | Coffee tweets: global + three brands | Topic modeling + importance–satisfaction opportunity mapping | Preference topics, opportunity topics, cross-brand recommendations | Connects social listening with decision-oriented coffee innovation and digital market sensing by transforming unstructured consumer discourse into structured opportunity priorities. |
Building on this gap, the study makes three contributions. First, it reconceptualizes social listening as a decision-oriented innovation process rather than a passive monitoring activity. Consumer conversations are systematically organized into preference structures and subsequently translated into directions for product development. Second, it combines a global coffee corpus with three brand-specific corpora, enabling a comparative examination of market-wide and brand-community opportunity structures. This design distinguishes broadly shared coffee-consumption themes from opportunities that arise within particular brand contexts. Third, it translates topic-level evidence into practical directions for product, service, promotional, and distribution innovation. These directions are intended to support coffee businesses with limited market research resources in prioritizing innovation hypotheses for subsequent testing and validation.
3. Methodology
The current study (Figure 1) employs a data-driven social media analytics approach for detecting customer preference structures and converting them into opportunity signals. The research design involves five phases: (1) online data acquisition, (2) text preprocessing, (3) customer preference topic modeling, (4) importance–satisfaction opportunity identification, and (5) cross-case analysis and recommendation development. The workflow is designed to make the transition from unstructured social-media discourse to managerial opportunity signals explicit and reproducible.

The empirical dataset consists of Twitter/X postings collected during 2023 and is organized into four corpora: a general coffee corpus and three brand-specific corpora for Kopi Kenangan, Starbucks, and Kopi Janji Jiwa. The general coffee corpus contains 172,017 retained messages. The brand corpora were constructed from brand-name queries (including spacing and hashtag variants of “Kopi Kenangan”, “Starbucks”, and “Kopi Janji Jiwa/Janji Jiwa”), whereas the general corpus used coffee-related terms such as “coffee” and “kopi”. English and Indonesian language posts were retained because the analysis combines a global category corpus with Indonesian brand communities.
Data collection covered the 2023 observation window. The archived research materials available for this revision do not preserve the day-level extraction log or the intermediate/final document counts for the three brand corpora; therefore, unsupported exact dates or tweet counts are not reconstructed. This provenance limitation is stated explicitly so that the reported sample size is not overstated. For future replication, the collection log should store query strings, start/end timestamps, raw counts, and post-cleaning counts for every corpus.
The use of two types of corpora is motivated by two analytical reasons. The general corpus allows capturing the discussion about coffee consumption in general without limiting the interpretation to one particular brand community. The brand corpora serve as a comparison case to understand the differences between the preference and opportunity structures in different market environments.
The empirical analysis comprised four analytical corpora collected during the 2023 observation window: a global coffee corpus and three brand-specific corpora representing Kopi Kenangan, Starbucks, and Kopi Janji Jiwa (Table 2). The global coffee corpus contained 172,017 retained tweets. The archived analytical outputs preserved the final preference-topic structures and importance–satisfaction opportunity maps for all four corpora; however, separate document counts for the three brand-specific corpora were not retained in the archived extraction records. Consequently, these counts are reported as unavailable rather than retrospectively reconstructed or estimated. This reporting decision was adopted to preserve the accuracy and traceability of the empirical data. Despite this limitation in corpus-size documentation, the analytical outputs were completely retained, comprising 13, 8, 12, and 15 preference topics for Global Coffee, Kopi Kenangan, Starbucks, and Kopi Janji Jiwa, respectively, from which 5, 3, 5, and 10 opportunity topics were identified.
Corpus | Period | Retained Tweet Count | Preference Topics | Opportunity Topics |
Global coffee | 2023 observation window | 172,017 | 13 | 5 |
Kopi Kenangan | 2023 observation window | — | 8 | 3 |
Starbucks | 2023 observation window | — | 12 | 5 |
Kopi Janji Jiwa | 2023 observation window | — | 15 | 10 |
Text preprocessing was performed before topic modeling to reduce noise while preserving preference-bearing expressions. Hyperlinks, user mentions, duplicated/reposted text, non-informative punctuation, and repeated whitespace were removed; text was lowercased and tokenized; standard English and Indonesian stop words were excluded; and obvious spelling/format variants were normalized. Reposts/retweets and exact duplicates were removed to avoid repeatedly weighting the same message. Posts dominated by promotional calls-to-action, repetitive advertising templates, or links without substantive consumer discussion were excluded, as were posts in which the query term was used in an irrelevant context. Brand-owned promotional content was treated conservatively and excluded when it functioned as advertising rather than consumer discourse. The retained English-and Indonesian-language text was then used for topic modeling.
In the third stage, latent themes within each corpus are modeled to represent consumer preference structures. Topics are seen as recurring groupings of semantically similar words and discussions that express an attribute of the product, consumption setting, service experience, behavior, or market concept. Such a view allows preference extraction to go beyond individual words and capture consumer concepts more broadly.
BTM was used as the final topic-modeling method. The model was implemented in Python using a standard BTM workflow with corpus-level biterm construction, symmetric topic priors (α = 50/K and β = 0.01), and 100 Gibbs-sampling iterations for each candidate solution, where K is the number of topics, and α and β are the symmetric Dirichlet priors for the topic distribution and the word distribution, respectively. Candidate topic numbers were screened over K = 5–20 for each corpus. The final K was selected by jointly considering topic coherence, stability of the main word groups across repeated runs, and managerial interpretability, rather than maximizing a single numerical criterion. This procedure resulted in K = 13 for the global coffee corpus, K = 8 for Kopi Kenangan, K = 12 for Starbucks, and K = 15 for Kopi Janji Jiwa. Topic labels were assigned after reviewing the highest-probability terms and their semantic relationship to product attributes, consumption situations, service experiences, and market mechanisms.
To improve interpretability, the Results section uses more specific descriptors for previously broad labels and distinguishes the two Janji Jiwa topping clusters as “Topping Variant I” and “Topping Variant II”. Because the archived files used for this revision do not contain the complete topic–word probability tables or original tweet-level examples, no top-word lists or quotations are fabricated; this limitation is explicitly acknowledged in Section 6.
The fourth step establishes a connection between the preference structure and product opportunity. For each topic k, importance was operationalized as normalized topic prevalence, i.e., the proportion of retained documents associated with the topic relative to the most prevalent topic in the same corpus:
$I_k=\frac{P_k}{\max (P)}$
Satisfaction was operationalized from the mean sentiment polarity of the documents associated with topic k and linearly rescaled to the 0–1 interval:
$S_k=\frac{s_k+1}{2}$
where, k indexes the topic, Ik is the normalized importance of topic k, Pk is the prevalence of topic k, max(P) is the maximum topic prevalence in the same corpus, Sk is the satisfaction score of topic k, and sk is the mean sentiment polarity of topic k, ranging from −1 to +1, so that Sk is rescaled to the 0–1 interval. Thus, both axes are dimensionless and comparable within each corpus. The reference crosshairs were set at 0.50 on both normalized axes, consistent with the maps reported in Figure 2, Figure 3, Figure 4, Figure 5. Topics with Ik ≥ 0.50 and Sk ≥ 0.50 were classified in the high-importance/high-satisfaction region.
In this study, the high-importance/high-satisfaction region is interpreted as a market-supported expansion opportunity rather than a corrective-improvement quadrant. A topic enters this region when it is both relatively salient and positively evaluated. It therefore signals an existing consumer-supported strength that can be reinforced through line extensions, complementary offerings, service enhancements, communication, or channel expansion. By contrast, a high-importance/low-satisfaction topic would be more naturally interpreted as a corrective-improvement priority.
Accordingly, the opportunity map is used as a screening device: it prioritizes positively supported themes for controlled market experimentation, while managerial validation remains necessary before commercialization.
The opportunity structures in the four sets of corpora are examined for the purpose of finding out three types of innovation signals, namely (1) market-level opportunities common to all brands, which are indicative of general consumption of coffee; (2) brand-specific opportunities based on the positioning or customer community of the brand; and (3) opportunity breadth, which is represented by the number of topics labeled as opportunities. Managerial implications are made through conversion of the themes identified into concrete actions in terms of products, services, promotions, and channels. It is important to note that these managerial implications are clearly stated as observations from social listening activities.
The framework provides analytical opportunity signals rather than evidence of commercial success. Validation is therefore positioned as a subsequent managerial stage: detected topics are translated into an innovation hypothesis, tested on a limited scale, and evaluated using observable outcomes such as sales, repeat purchase, attachment rate, engagement, or customer feedback. The present study does not claim that these market tests were empirically conducted; instead, Section 4.6 demonstrates how the identified topics can be converted into testable managerial actions.
4. Results and Discussion
The analysis revealed different levels of innovation-related insight across the four social media corpora. The global coffee corpus primarily captured broad consumption contexts, functional benefits, and experiential meanings, whereas the brand-specific corpora revealed more concrete combinations of product attributes, service experiences, pricing considerations, promotional mechanisms, customer segments, and distribution channels. The following subsections present the preference structures and market-supported opportunity spaces identified for the global coffee corpus, Kopi Kenangan, Starbucks, and Kopi Janji Jiwa. The cross-corpus comparison subsequently examines differences in the composition and breadth of these opportunity structures without treating them as direct measures of brand performance or innovation capability.
The global coffee corpus comprised 172,017 tweets and yielded 13 consumer preference topics. Five topics were located in the high-importance and high-satisfaction region and were therefore classified as market-supported opportunity topics: Topic 2, Benefits of Coffee for Work; Topic 3, Coffee Shop; Topic 5, Enjoyable Drink; Topic 9, Caffeine and Sleep; and Topic 11, Coffee in the Morning. Accordingly, five of the 13 modeled topics, representing 38.5% of the topic structure, met the within-corpus opportunity-screening criterion.
For interpretive clarity, Enjoyable Drink refers to a hedonic consumption theme in which coffee was discussed in relation to pleasure and enjoyment rather than a specific functional benefit. It should not be interpreted merely as a generic positive-sentiment category. Similarly, Caffeine and Sleep represents discourse concerning the relationship between caffeine consumption, alertness, and sleep-related considerations rather than an unqualified endorsement of increased caffeine intake.
As shown in Figure 2, the five topics represent several dimensions of coffee-related value. Benefits of Coffee for Work reflects a functional dimension; Coffee Shop captures the role of place and service experience; Enjoyable Drink represents hedonic value; Caffeine and Sleep concerns physiological and temporal considerations; and Coffee in the Morning reflects routine-based consumption. The resulting opportunity structure indicates that coffee innovation extends beyond beverage formulation. It encompasses the interaction among product design, consumption occasion, functional value, service environment, customer experience, and market communication.

The findings were consistent with previous social media research showing that consumers associate coffee with energy, well-being, mood, and lifestyle (Samoggia et al., 2020). The present analysis extends this line of research by positioning these consumer meanings within an importance–satisfaction opportunity framework. From a managerial perspective, innovation screening should therefore consider not only flavor development but also work-related consumption, morning routines, coffee-shop experiences, and responsible communication concerning caffeine use. These results identify possible directions for controlled market experimentation rather than confirming the commercial viability of any specific product or intervention.
The Kopi Kenangan corpus yielded eight consumer preference topics, three of which were classified as opportunity topics. These topics represented 37.5% of the modeled preference structure and comprised Topic 4, In-Shop Service; Topic 6, The Bittersweet Taste of Coffee; and Topic 8, Milk Variant. Their positions within the high-importance and high-satisfaction region are presented in Figure 3.

Compared with the broader opportunity structure of the global corpus, the Kopi Kenangan structure was more concentrated around the connection between the core beverage and the point-of-service experience. In-Shop Service represents the service-delivery dimension, The Bittersweet Taste of Coffee reflects sensory positioning, and Milk Variant indicates flexibility in beverage configuration. Together, these topics suggest a focused innovation pathway in which the established sensory identity of the product can be reinforced while selected service processes and milk-based variants are tested.
This relatively concentrated structure may be particularly useful for resource-constrained product development. Rather than pursuing extensive and simultaneous diversification, a coffee business could use the identified themes to prioritize a small number of controlled interventions. Possible applications include testing selected service routines, adjusting the balance of bitter and sweet sensory characteristics, or introducing a limited number of milk-based variants aligned with the existing taste profile. The topics provide market-supported directions for experimentation, but further validation through sales, repeat purchase, and direct customer feedback remains necessary.
The Starbucks corpus yielded 12 consumer preference topics, five of which entered the opportunity region. These five topics represented 41.7% of the modeled topic structure: Topic 1, Place to Drink; Topic 4, Price and Taste; Topic 7, Recipe and Snack; Topic 8, Merchandise; and Topic 12, Kid-Friendly Starbucks Drinks. Their positions in the high-importance and high-satisfaction region are shown in Figure 4.

The label Place to Drink represents an outlet and consumption-space theme. It captures the role of Starbucks as a physical setting in which customers consume beverages, spend time, and engage in individual or social activities. It is therefore analytically distinct from beverage attributes such as taste, price, and recipe.
The Starbucks opportunity structure extended beyond the core beverage. Place to Drink reflected the spatial and experiential function of the outlet; Price and Taste combined value assessment with product evaluation; Recipe and Snack indicated complementary consumption; Merchandise represented a non-beverage extension of the brand; and Kid-Friendly Starbucks Drinks indicated the potential relevance of a distinct customer segment. Together, these themes formed a product–experience ecosystem encompassing the beverage, the consumption environment, complementary products, brand extensions, and segment-specific offerings.
These findings have practical relevance beyond Starbucks itself. Smaller coffee businesses do not need to reproduce the complete product portfolio or market architecture of a large international chain. Instead, they can identify adjacent dimensions of value that are compatible with their positioning and operational capacity. A resource-constrained business might focus on improving the consumption-space experience or developing a carefully selected beverage–snack pairing rather than simultaneously expanding its product range, customer segments, and merchandise. In this context, social listening provides evidence for selecting feasible innovation hypotheses, while controlled testing remains necessary before wider implementation.
The Kopi Janji Jiwa corpus yielded 15 consumer preference topics, ten of which were located in the high-importance and high-satisfaction region. These ten topics represented 66.7% of the modeled topic structure: Topic 1, Food and Snack; Topic 2, Topping Variant I; Topic 3, Discount and Promotion; Topic 4, Friend Recommendations; Topic 5, Topping Variant II; Topic 7, Milk and Tea Variant; Topic 9, Advertising; Topic 12, Place and Shop; Topic 13, Appeal to Millennials; and Topic 15, Channels and Collaborations (Figure 5).

The two topping-related clusters were retained as separate topics because they emerged as distinct latent word groups in the original modeling results. The Roman numerals distinguish the two clusters and prevent them from being interpreted as accidental duplication where the complete topic–word probability tables are unavailable.
Compared with the other corpora, the Kopi Janji Jiwa opportunity structure covered a broader range of product-architecture and market-activation mechanisms. Food and Snack, Topping Variant I, Topping Variant II, and Milk and Tea Variant represented menu-related innovation. Discount and Promotion, Advertising, and Channels and Collaborations reflected mechanisms for market development. Friend Recommendations and Appeal to Millennials captured social influence and segment positioning, while Place and Shop represented the experiential role of the outlet.
This thematic breadth indicates that consumer discourse concerns not only what is consumed but also how offerings are combined, promoted, recommended, accessed, and experienced. From an innovation-management perspective, product performance may depend on the alignment of beverage or food development with communication, social diffusion, customer segmentation, and channel design. A new topping or beverage variant, for example, may require appropriate promotional support, recommendation mechanisms, and distribution or collaboration channels. The resulting opportunity structure therefore provides a basis for integrated product–market experimentation. However, the larger opportunity-topic share should not be interpreted as evidence that Kopi Janji Jiwa has greater innovation capability or market potential than the other brands.
Table 3 summarizes the preference topics and market-supported opportunity topics identified across the four corpora. The comparison is descriptive and focuses on differences in the composition and breadth of the themes that met the within-corpus opportunity criterion. The global coffee corpus emphasized consumption occasions, functional benefits, and experiential meanings, whereas the brand-specific corpora revealed more concrete combinations of products, services, places, prices, promotions, customer segments, and channels.
Corpus | Preference Topics | Opportunity Topics | Opportunity-Topic Share | Reported Opportunity Themes |
Global coffee | 13 | 5 | 38.50% | Benefits of Coffee for Work; Coffee Shop; Enjoyable Drink; Caffeine and Sleep; Coffee in the Morning |
Kopi Kenangan | 8 | 3 | 37.50% | In-Shop Service; The Bitter Sweet Taste of Coffee; Milk Variant |
Starbucks | 12 | 5 | 41.70% | Place to drink; Price & Taste; Recipe & Snack; Merchandise; Kid-friendly Starbucks drinks |
Kopi Janji Jiwa | 15 | 10 | 66.70% | Food & Snack; Topping Variant I; Discount & Promotion; Friend Recommendations; Topping Variant II; Milk & Tea Variant; Advertising; Place & Shop; Appeal to Millennials; Channels & Collaborations |
The opportunity-topic shares of 38.5%, 37.5%, 41.7%, and 66.7% indicate the proportion of modeled topics that entered the high-importance and high-satisfaction region within each corpus. They are not inferential statistics and should not be treated as directly comparable measures of brand-level innovation potential.
The opportunity-topic shares require cautious interpretation because the four corpora differed in size and the selected number of topics was specific to each corpus. A higher share does not demonstrate that one brand possesses stronger innovation capability, greater market potential, or superior business performance. It indicates only that a larger proportion of the topics generated for that corpus met the study’s within-corpus importance–satisfaction criterion. The cross-brand analysis therefore concerns the composition and breadth of the detected themes rather than a ranking of the brands.
The findings further indicate that a universal approach to coffee innovation would be inappropriate. In the global corpus, the opportunity structure centered on consumption routines and the functional, experiential, and hedonic meanings of coffee. For Kopi Kenangan, the opportunity themes were concentrated around in-shop service, sensory positioning, and beverage variants. The Starbucks structure connected place, value assessment, complementary products, merchandise, and customer segments. For Kopi Janji Jiwa, product variants were accompanied by promotions, recommendations, advertising, experiential settings, and collaboration channels. These differences should be interpreted as corpus-specific opportunity configurations rather than evaluations of the innovation capacity of individual brands. The same social listening framework can therefore generate different innovation portfolios depending on the consumer community and market context being examined.
This result is consistent with the broader innovation-management view that social media analytics can provide external market intelligence that is difficult to obtain through surveys based exclusively on predefined attributes (Geissinger et al., 2023). It also aligns with automated marketing research showing that naturally occurring online customer reviews can reveal consumer-relevant product attributes without relying solely on predetermined survey categories (Lee & Bradlow, 2011).
Opportunity Domain | Evidence from the Study | Possible Innovation Action | Small Market Test | Coffee Business Decision Logic |
Consumption occasion and functional value | Work benefits; morning coffee; caffeine & sleep | Develop offers or communication tailored to time/use occasions; clarify caffeine positioning | Pilot an occasion-based bundle or communication campaign for a limited period or selected outlets. | Use low-cost occasion-based bundles or campaigns before investing in new formulations. |
Sensory and beverage variants | Bittersweet taste; milk variant; topping variants; milk & tea variant | Refine taste profiles and expand variants closely related to proven preferences | Introduce one or two selected variants in a limited number of outlets or for a limited period rather than immediately expanding the full product line. | Prioritize modular ingredients that allow several variants from a limited inventory base. |
Experience and place | Coffee Shop; In-Shop Service; Place to drink; Place & Shop | Improve service flow, ambience, and the role of the outlet as a consumption space | Test a specific service or outlet-experience change in selected outlets before wider implementation. | Test operational/service changes that improve experience without major capital expenditure. |
Complementary products | Recipe & Snack; Food & Snack; merchandise | Build selective add-on products or pairings around the core beverage | Offer a limited selection of snack or complementary-product pairings with selected core beverages. | Choose complements with clear cross-selling potential and manageable complexity. |
Promotion and social influence | Discount & Promotion; Friend Recommendations; Advertising; millennial appeal | Use targeted promotion, referral mechanisms, and segment-specific communication | Run a targeted promotion or referral campaign for a defined customer segment or limited campaign period. | Link campaigns to measurable conversion, repeat purchase, or referral outcomes. |
Channels and collaboration | Channels & Collaborations | Develop delivery, partnership, co-branding, or community collaboration channels | Pilot one delivery, partnership, co-branding, or community collaboration initiative with a limited scope. | Use partnerships to access capabilities or audiences that are costly to build internally. |
The managerial value of social listening becomes clearer when the identified opportunity topics are translated into testable categories of action. Table 4 connects the empirical opportunity themes with potential innovation responses, limited-scale market tests, and decision principles relevant to coffee businesses. These proposed actions are not evidence of implementation or guarantees of commercial success. They represent managerial hypotheses derived from the analytical results and should be evaluated through sequential market testing.
Two examples illustrate how the framework can be applied. For Kopi Kenangan, Milk Variant should not be interpreted as proof that any milk-based product will succeed. Instead, the topic can be converted into a constrained market experiment in which one or two milk-based recipes are introduced in selected outlets or during a limited period. Consumer responses can then be evaluated through sales, repeat purchase, and direct feedback before a decision is made regarding modification, expansion, or discontinuation.
Similarly, Recipe and Snack in the Starbucks corpus can be translated into a limited test of selected beverage–snack combinations. The response can be assessed through attachment rates, incremental sales, repeat purchase, and customer feedback. This staged interpretation preserves the decision relevance of social listening while avoiding the assumption that online discussion alone establishes demand.
Social listening should therefore be understood as an innovation-prioritization mechanism rather than an automatic product-generation tool. Its managerial application involves four sequential stages: identifying a topic-level signal, formulating an innovation hypothesis, conducting a controlled market test, and evaluating the resulting sales or consumer-response evidence. This staged process is particularly relevant to resource-constrained coffee businesses because it preserves the speed and scalability of digital market sensing while reducing the risk of investing in themes that merely reflect temporary social media attention. A favorable market response can justify refinement or expansion, whereas a weak response indicates that the proposed intervention should be modified, retested, or discontinued.
The findings have three implications for data-driven innovation research. First, the study positions external digital knowledge as an input to product opportunity identification. Unlike approaches limited to sentiment classification or descriptive summaries of online discourse, the proposed framework connects consumer preference structures with opportunity mapping and corresponding categories of managerial action. It therefore establishes a more explicit analytical path from digital consumer expression to innovation prioritization.
Second, the cross-corpus design demonstrates that innovation signals are context-dependent. The global corpus captured broad meanings and occasions associated with coffee consumption, whereas the brand-specific corpora revealed different combinations of product, service, experience, promotion, segment, and channel mechanisms. Future social media analytics research should therefore distinguish between category-level and brand-specific signals rather than assume that a single corpus can represent the complete opportunity environment.
Third, the findings reinforce the strategic relevance of social media for coffee businesses operating under resource constraints. Previous research indicates that social media use can support innovation capability and firm performance (Borah et al., 2022; Noviaristanti et al., 2023). The present study extends this argument by demonstrating how social media use can be operationalized as an analytical procedure for detecting and prioritizing innovation opportunities. Its contribution lies not in claiming that digitally identified topics guarantee market success, but in showing how they can be converted into structured hypotheses for subsequent business validation.
5. Conclusion
This study developed a product innovation framework that combined social listening, consumer preference modeling, and importance–satisfaction opportunity identification. The empirical analysis examined four Twitter/X corpora collected during 2023: a global coffee corpus comprising 172,017 tweets and three brand-specific corpora associated with Kopi Kenangan, Starbucks, and Kopi Janji Jiwa. The framework transformed unstructured consumer discourse into preference topics, identified market-supported opportunity topics, compared their composition across the four corpora, and translated the resulting themes into testable innovation directions.
The global coffee corpus yielded 13 preference topics, five of which represented opportunities associated with work-related benefits, coffee-shop experiences, drinking enjoyment, the relationship between caffeine and sleep, and morning consumption. The Kopi Kenangan corpus yielded eight preference topics and three opportunity topics concerning in-shop service, bittersweet taste, and milk variants. The Starbucks corpus produced 12 preference topics and five opportunity topics relating to place, price and taste, beverage–snack combinations, merchandise, and child-oriented offerings. The Kopi Janji Jiwa corpus yielded 15 preference topics and ten opportunity topics spanning food and snacks, topping variants, milk and tea variants, promotions, recommendations, advertising, place, millennial appeal, and collaboration channels.
These findings demonstrate that coffee innovation cannot be confined to beverage attributes. The identified opportunities encompass consumption occasions, functional benefits, sensory characteristics, service experiences, complementary products, promotional mechanisms, social influence, customer segments, distribution channels, and collaborative arrangements. For resource-constrained coffee businesses, the framework provides a practical means of reducing large volumes of online discourse to a manageable set of innovation hypotheses that can be evaluated before substantial resources are committed.
Methodologically, the study shows the value of connecting topic identification with decision-oriented opportunity assessment. The proposed framework advances data-driven innovation research by transforming naturally occurring consumer discourse into structured opportunity priorities while preserving the need for managerial interpretation and market validation. Social listening therefore functions as a technology-enabled market-sensing and innovation-prioritization mechanism rather than as a substitute for product testing or evidence of commercial success.
6. Limitations and Future Research
This study has several limitations. First, the empirical evidence was derived exclusively from Twitter/X. Consumer discourse, platform demographics, communication practices, and recommendation mechanisms differ across digital environments and may also change over time. Future research should determine whether the identified opportunity structures remain stable when more recent data and additional sources, such as Instagram, TikTok, review platforms, online communities, or food-delivery applications, are analyzed.
Second, social media users do not constitute a statistically representative sample of coffee consumers. The findings should therefore be interpreted as digital market signals rather than population-level estimates of consumer preferences. Subsequent studies could combine social listening with surveys, interviews, controlled experiments, transaction records, loyalty data, or sales information. Such triangulation would strengthen external validity and make it possible to examine whether topic-level opportunity signals predict observable consumer behavior.
Third, the number and interpretation of the identified topics depended on the selected modeling configuration and the availability of topic-level outputs. The archived materials available for the present revision did not preserve the complete topic–word probability tables, representative tweet-level examples, day-level extraction records, or post-cleaning document counts for the three brand-specific corpora. This limitation reduces the degree to which the reported topic structures can be independently reproduced and should be considered when interpreting the cross-brand findings. Future replication packages should retain complete query strings, extraction dates, raw and cleaned corpus counts, preprocessing records, topic–word distributions, topic assignments, model-selection statistics, and anonymized representative texts.
Future research should also compare BTM with LDA, BERTopic (a transformer-based topic modeling approach), and other short-text or embedding-based topic models. Such comparisons should evaluate topic coherence, stability across repeated runs, sensitivity to the selected number of topics, and agreement among domain experts. In addition, aspect-level sentiment analysis could provide a more precise satisfaction measure when a single post discusses several product or service attributes with different evaluative orientations. Further validation through longitudinal data and controlled market experiments would help determine whether opportunity topics identified through social listening led to measurable improvements in product adoption, repeat purchase, customer engagement, or business performance.
The data used to support the research findings are available from the corresponding author upon request.
The author declares no conflicts of interest.
