A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success
Abstract:
Social media provides a rich source of consumer-generated data that can support product innovation decisions. However, firms still face difficulties in transforming unstructured online discussions into measurable insights that are linked to market success. This study proposes a data-driven social media framework for product innovation by integrating text mining, Principal Component Analysis (PCA), and Principal Component Regression (PCR). The framework identifies dominant consumer attribute discussions, reduces correlated attributes into interpretable latent components, and tests their relationship with standardized sales indicators. The study uses two competitive smartphone products, Apple iPhone 8 and Samsung Galaxy Note 7, as empirical cases. Text mining results show that consumers discussed visual, ecosystem, connectivity, battery, charging, camera, and platform-related attributes across both products. PCA results indicate that a small number of principal components can explain most of the variance in consumer discussions. For the iPhone 8, the first two components explain 79% of the total variance, while for the Galaxy Note 7, they explain 81%. The regression results show strong links between selected social media-derived components and market success. For the iPhone 8, PC1 has the strongest relationship with standardized sales, with an R² of 0.956. For the Galaxy Note 7, PC7 shows the strongest relationship, with an R² of 0.889, while other components show both positive and negative relationships. These findings show that social media discussions can provide early signals of consumer needs, product risks, and market response. The proposed framework offers a systematic tool for attribute prioritization, launch monitoring, and evidence-based product development.1. Introduction
Product innovation is essential for firms that compete in dynamic and technology-driven markets. Rapid changes in customer preferences, shorter product life cycles, and intense competition require firms to develop products that match market expectations. Product success depends not only on technological novelty, but also on the ability of the innovation to diffuse and gain acceptance in the market (Klein & Şener, 2023). Therefore, firms need to understand customer needs accurately and use this understanding as the basis for product development decisions (Lin, 2018; Xie & Jia, 2016).
Despite its importance, product innovation remains risky. Many product ideas fail before they reach commercial success. Purnama et al. (2023) reported that less than one percent of initial product development ideas are commercialized. Cooper (2019) also showed that only a small proportion of product concepts that enter the development and testing process become commercially successful. These failures are often linked to changing market needs, high research and development costs, limited resources, and shorter product life cycles (Najafi-Tavani et al., 2020). These conditions show that firms need more effective methods to identify customer needs at the early stage of new product development (NPD).
Customer need identification plays a central role in NPD. Ulrich et al. (2008) emphasized that a comprehensive understanding of customer needs should begin in the early development stage, including data collection, need interpretation, importance assessment, and opportunity identification. Customer involvement also supports product innovation because it helps firms reduce the risk of developing products that do not fit market expectations (Cui & Wu, 2016; Harvey & Kou, 2013; Morgan et al., 2018). This perspective is consistent with recent studies on user innovation and value co-creation, which view customers as active contributors to idea generation, product design, and innovation development. In NPD, user involvement can help firms identify relevant needs, refine product concepts, and reduce uncertainty during early development stages. However, traditional customer need identification methods, such as surveys, interviews, and focus groups, remain time-consuming and resource-intensive because they depend on structured data collection, participant recruitment, and manual interpretation (Nasrabadi et al., 2024; Sońta-Drączkowska et al., 2025). Recent studies show that user-generated content and social media analytics offer scalable alternatives for capturing customer knowledge from online platforms and transforming it into innovation insights (Geissinger et al., 2023; Nasrabadi et al., 2024). Therefore, firms need faster, more scalable, and more cost-efficient approaches to support customer need identification and data-driven product planning (Panzner et al., 2024; Tian et al., 2024).
Social media offers a promising source of customer insight for product development. Consumers use social media to express expectations, preferences, complaints, and comparisons between competing products. These discussions provide real-time information about product attributes, perceived value, and potential adoption barriers. Recent literature confirms that social media analytics has become an important methodological tool in innovation management because it can capture and analyze user-generated content from online platforms (Geissinger et al., 2023). Social media also supports different phases of NPD, including discovery, development, and launch monitoring (Han et al., 2025). In this context, online consumer discussions can help firms identify emerging needs, detect product issues, and evaluate early market responses.
User-generated content has also received growing attention in NPD research. Nasrabadi et al. (2024) reviewed studies on user-generated content in NPD and identified four major themes, namely its impact on NPD, idea mining, feature derivation, and customer need understanding. This shows that online consumer data is not only useful for marketing analysis, but also for product design and innovation decisions. However, social media data is unstructured, informal, noisy, and difficult to interpret directly. Firms need a systematic method to transform online consumer discourse into measurable and actionable product development insights.
Previous studies have used social media analytics to support product development. Zhang et al. (2018) identified important topics using clustering methods. Hong & Wang (2021) summarized customer opinions to identify product attributes. Other studies applied sentiment analysis to measure satisfaction and dissatisfaction in online discussions (Rathore & Ilavarasan, 2020; Zhang et al., 2018) . Recent studies also show that social media analytics can help firms mine customer complaints and identify product opportunities. Wang et al. (2023) developed a social media analytics method that combines sentiment analysis, topic modeling, topic engagement, and topic emergence to explore product improvement opportunities from customer complaints. Similarly, Zhang & Song (2024) showed that product improvement decisions can be strengthened by combining consumer big data, text mining, and group decision-making methods.
Although these studies show the value of social media for product development, several gaps remain. Many studies focus on topic discovery, sentiment classification, or customer complaint mining. These approaches are useful, but they often remain descriptive. They do not fully explain how product attributes discussed on social media form latent consumer need structures. They also do not sufficiently test how those structures relate to measurable product success. Recent literature on data-driven product planning also stresses the need for transparent analytics pipelines that connect data sources, preprocessing, modeling, and evaluation in a systematic way (Panzner et al., 2024). This gap is important because firms need not only to know what consumers discuss, but also to identify which dimensions of discussion are most related to market outcomes.
This study addresses that gap by proposing a data-driven social media framework for product development. The framework integrates text mining, Principal Component Analysis (PCA), and Principal Component Regression (PCR) to connect consumer attribute discussions with market success. Text mining is used to extract dominant product attributes from social media discussions. PCA reduces correlated attributes into a smaller set of interpretable latent components. The PCR then examines the relationship between these components and a quantitative product success indicator.
This study uses two competitive smartphone products, Apple iPhone 8 and Samsung Galaxy Note 7, as empirical cases. These products provide a relevant context because smartphone markets are highly competitive, attribute-driven, and strongly influenced by consumer discourse. By comparing two competing products, this study shows how different patterns of consumer attribute discussions can be linked to different market outcomes.
This study makes three main contributions. First, it develops a systematic framework that transforms unstructured social media text into measurable product development insights. Second, it applies PCA to identify latent structures among consumer-discussed product attributes, which helps reduce redundancy and multicollinearity in textual attribute data. Third, it links these latent components to product success through regression analysis, allowing firms to identify which dimensions of consumer discourse are most relevant to market performance. The proposed framework can support attribute prioritization, launch monitoring, competitive benchmarking, and early risk detection in product development.
2. Literature Reviews
Customer involvement plays an important role in NPD because it helps firms understand market needs, refine product concepts, and reduce uncertainty in product decisions. Morgan et al. (2018) found that customer participation is positively related to NPD performance and that this relationship is mediated by innovativeness. This finding shows that customers do not only act as product users, but also as external knowledge sources that can support product innovation (Morgan et al., 2018).
Recent research also shows that user-generated content has become a relevant data source for NPD. Nasrabadi et al. (2024) reviewed user-generated content studies in NPD and identified several key themes, including idea mining, feature derivation, customer need understanding, and the impact of online user content on product development. These themes show that online consumer data can help firms identify product attributes, detect emerging needs, and support product improvement decisions.
In this context, social media expands the role of customer involvement. It allows firms to observe consumer discussions in real time and at a larger scale than traditional customer research methods. Han et al. (2025) reviewed 110 studies from 2002 to 2023 and showed that social media can support three phases of NPD: discovery, development, and launch. Their study also identified several NPD objectives that can be addressed through social media, including idea generation, customer need identification, product evaluation, and launch monitoring.
Social media analytics has emerged as an important approach in innovation management and product development. It allows firms to capture and analyze user-generated content from online platforms. Geissinger et al. (2023) stated that social media analytics can support customer-focused, market-focused, technology-focused, and society-focused innovation research. Their systematic review also shows that the use of social media analytics in innovation management has grown during the last decade, although the field is still developing.
Social media is useful for product development because consumers often discuss product attributes, usage problems, complaints, comparisons, and expectations on online platforms. These discussions can reveal product strengths and weaknesses that may not be captured through formal surveys or interviews. In business-to-business (B2B) contexts, Tian et al. (2024) showed that strategic use of social media can help firms acquire external knowledge, strengthen absorptive capacity, and improve NPD. This supports the view that social media can function as a knowledge source for product planning and innovation decisions.
However, social media data is difficult to use directly. It is unstructured, informal, noisy, and highly dynamic. This creates a methodological challenge for firms that want to convert online discussions into product development insights. Therefore, a systematic analytics framework is needed to process text data, extract relevant product attributes, reduce data complexity, and connect consumer discourse with measurable product outcomes.
Text mining is widely used to extract meaningful information from unstructured consumer data. In product development research, text mining helps identify product attributes, customer opinions, usage problems, and potential improvement opportunities. Zhang et al. (2018) used online review data mining to support product innovation and improvement. Their study shows that online consumer reviews can reveal product-related issues and support product design decisions.
Hong & Wang (2021) developed a customer opinion summarization approach using topic mining and deep neural networks. Their study highlights the importance of extracting product attributes and related opinions from large volumes of customer text. This approach helps firms summarize customer perceptions and identify attribute-level insights that can guide product improvement.
Rathore & Ilavarasan (2020) applied Twitter/X analytics to analyze pre-launch and post-launch emotions in NPD. Their study shows that social media can capture changes in consumer emotion across product launch stages. This is important because consumer reactions before and after launch can indicate acceptance, dissatisfaction, or potential product risk.
Recent studies have moved beyond simple text extraction by combining text mining with other analytical methods. Wang et al. (2023) developed a social media analytics method to mine customer complaints and identify product opportunities. Their method combines preprocessing, sentiment analysis, topic modeling, topic engagement, and opportunity evaluation. This shows that text mining can support product opportunity discovery when it is combined with structured analytical procedures.
Zhang & Song (2024) also proposed a product improvement method in a big data environment by combining text mining and large group decision-making. Their findings show that text mining can improve product improvement decisions when it is integrated with decision-support methods. This confirms that customer text data needs to be transformed into structured knowledge before it can guide product development actions.
Data-driven product planning requires firms to connect customer data, analytical methods, and product decisions through a clear process. Panzner et al. (2024) argued that successful data-driven product planning requires a structured pipeline that links domain knowledge, data analysis, and product planning tasks. Their study proposes a data analytics toolbox to support product planning, showing the need for systematic methods that can guide firms from raw data to actionable product decisions. This view is also consistent with recent studies on online review-based decision support, which show that customer-generated data can be transformed into attribute-level preference analysis and product ranking models (Cai & Zhang, 2025; Wu et al., 2024).
One major issue in social media-based product analysis is the presence of many correlated product attributes. Consumers often discuss attributes together, such as display and camera, battery and charging, or operating system and ecosystem compatibility. If these variables are analyzed separately, the model may suffer from redundancy and multicollinearity. Multicollinearity can reduce model stability, inflate estimation uncertainty, and weaken the interpretability of individual predictors (Greenacre et al., 2022; Hair et al., 2019). In product development research, this issue is important because consumer needs are often multidimensional and interrelated, so single-attribute analysis may fail to capture the broader structure of customer perception (Wu et al., 2024; Zhang & Song, 2024).
The PCA can address this issue by reducing correlated attributes into a smaller set of independent components. PCA transforms a cases-by-variables data table into a smaller number of principal components that retain the main information in the original variables while improving interpretability (Greenacre et al., 2022). Each principal component represents a latent structure in consumer discussions. In product development research, this is useful because consumer needs are often multidimensional and cannot be fully represented by single attributes. PCA helps identify broader dimensions of customer perception, such as visual experience, ecosystem compatibility, charging performance, design preference, or platform identity. Therefore, PCA provides a suitable dimensionality reduction method for converting complex social media-based product attributes into interpretable latent components before further regression analysis.
Many previous studies use social media analytics to identify product topics, sentiments, complaints, or customer needs. These studies provide valuable descriptive insights for product development, especially in identifying product opportunities, customer complaints, and improvement priorities (Wang et al., 2023; Wu et al., 2024). However, fewer studies examine how social media-derived attributes relate to market success.
This creates an important research gap. Firms need to know not only what consumers discuss, but also which dimensions of discussion are associated with product performance. Prior studies on online reviews show that consumer-generated text can support product sales forecasting and attribute importance prediction, indicating that online discourse can contain signals related to market response (Ding et al., 2025; Yakubu & Kwong, 2021).
The PCR can help address this gap. PCR combines PCA and regression analysis by first transforming correlated predictors into principal components and then using selected components as predictors in a regression model (Greenacre et al., 2022). This approach is suitable for social media-based product development because consumer-discussed attributes are often correlated. PCA reduces these correlated attributes into independent components, while regression tests the relationship between these components and a product success indicator. Therefore, PCR can reduce multicollinearity and help identify which latent dimensions of consumer discourse have stronger links to market outcomes.
In this study, market success is represented through a standardized sales indicator. This allows the study to examine whether consumer attribute discussions on social media can explain variation in product success. By applying PCR, the framework moves beyond descriptive text mining and provides a predictive link between online discourse and product performance. This position is consistent with recent data-driven product development studies, which emphasize the need to transform customer-generated data into measurable decision-support variables for product improvement and market-oriented decision-making (Panzner et al., 2024; Zhang & Song, 2024)
Based on the previous studies and Table 1, social media analytics has been widely used to support product development, especially for identifying product topics, customer needs, complaints, emotions, and improvement opportunities. Most studies apply text mining, topic modeling, sentiment analysis, or customer complaint mining to extract insights from online consumer discussions. These approaches are useful because they help firms understand what consumers discuss and how consumers respond to product attributes.
Author(s) | Study Case | Data Type | Data Source | Text | Advanced Modeling | Research Focus |
Zhang et al. (2018) | Huawei | UGC | OR | TM | – | PD |
Rathore & Ilavarasan (2020) | 3 products | UGC | SM | TM + SA | – | MS + PD |
Hong & Wan (2021) | Product reviews | UGC | OR | TM + Topic | – | PD |
Yakubu & Kwong (2021) | Reviews + Trends | UGC | OR | TM | Reg | PD |
Geissinger et al. (2023) | Innovation mgmt SLR | SLR + UGC | SM | – | – | PD |
Wang et al. (2023) | Complaints mining | UGC | SM | TM + Topic + SA | – | PD |
Panzner et al. (2024) | Data toolbox SLR | SLR + UGC | – | – | – | PD |
Nasrabadi et al. (2024) | UGC in NPD SLR | SLR + UGC | SM+OR | – | – | PD |
Tian et al. (2024) | B2B SM use | UGC | SM | – | – | PD |
Zhang & Song (2024) | Big data improv. | UGC | OR | TM | – | PD |
Han et al. (2025) | SM across NPD SLR | SLR + UGC | SM | – | – | PD |
Ding et al. (2025) | NEV sales forecast | UGC | OR | TM + SA | Reg | MS + PD |
This study | iPhone 8 & Note7 | UGC | SM | TM | PCA + Reg | MS + PD |
SM = Social Media; OR = Online Review; MS = Market Success; PD = Product Development; SLR = Systematic Literature Review;
UGC = user-generated content; en dash “–” indicates the corresponding method or focus was not covered in the cited study.
However, three important gaps remain. First, many studies still focus on descriptive analysis. They identify topics, sentiments, or complaints, but do not explain the latent structure among product attributes discussed by consumers. In social media discourse, product attributes are often correlated. For example, consumers may discuss display, camera, battery, operating system, and charging features together. If these attributes are analyzed separately, the results may contain redundancy and multicollinearity.
Second, prior studies rarely connect social media-derived product attributes with measurable market success. Most existing studies stop at identifying customer opinions or product improvement opportunities. They do not test whether the extracted attributes are related to product performance indicators, such as sales or market response. This limits the practical value of social media analytics for product development decisions.
Third, there is still limited integration between text mining, dimensionality reduction, and regression analysis in one systematic product development framework. Existing approaches often use text mining or sentiment analysis as standalone tools. A more integrated method is needed to transform unstructured social media text into latent consumer need components and then examine their relationship with product success.
Therefore, this study addresses these gaps by proposing a data-driven social media framework for product development that integrates text mining, PCA and PCR. The framework identifies consumer attribute discussions, reduces correlated attributes into interpretable latent components, and links these components to market success. The proposed framework differs from previous studies in social media analytics in three main aspects. Methodologically, it combines text mining, PCA and PCR into one single pipeline, rather than applying each method as a separate tool. Previous studies have typically used text mining or topic modelling to identify product attributes or consumer sentiments without examining how those attributes form latent structures or how those structures relate to quantifiable market outcomes. The research extends beyond providing descriptive insights at the application level. Instead of only identifying what consumers discuss, the framework examines which latent discussion dimensions are most strongly associated with market success, distinguishing between attributes that signal product appeal and those that indicate product risk. This in turn makes the framework suitable for attribute discovery, launch monitoring, competitive benchmarking and early risk detection in product development.
3. Methodology
This study develops a data-driven social media framework for product development by linking consumer attribute discussions with market success. The framework uses social media text as the main empirical input because online consumer discussions contain direct expressions of product expectations, complaints, comparisons, and perceived product attributes. The method combines text mining, PCA and PCR to transform unstructured consumer discourse into measurable product development insights.
The proposed framework consists of five main phases (Figure 1). The first phase collects and cleans social media data related to the selected products. The second phase preprocesses the text data to prepare it for analysis. The third phase extracts product attributes and builds an attribute frequency matrix. The fourth phase applies PCA to reduce correlated attributes into latent consumer need components. The fifth phase uses PCR to examine the relationship between these components and market success. The framework is applied to two competitive smartphone products, Apple iPhone 8 and Samsung Galaxy Note 7.

The first phase collected consumer discussions from social media related to Apple iPhone 8 and Samsung Galaxy Note 7. These two products were selected based on three historical considerations. First, both products represent flagship smartphones from two major competing ecosystems, Apple iOS and Samsung Android, which often generate intense consumer comparisons in terms of display, camera, battery, operating system, charging, design, and brand preference. Second, both products had strong public visibility during their launch periods, making them suitable for identifying consumer comments from online discussions. Third, the two products provide contrasting historical market outcomes. Therefore, the comparison between iPhone 8 and Samsung Galaxy Note 7 provides a relevant empirical setting to examine how different patterns of consumer attribute discussions relate to different market success outcomes.
Data collection focused on posts that contained product names, competing brand terms, and smartphone-related attributes. The search keywords included product-related terms such as “iPhone 8”, “Apple iPhone 8”, “Samsung Galaxy Note 7”, “Galaxy Note 7”.
To improve data quality, several cleaning criteria were applied. Posts were retained when they mentioned the selected products and contained discussions related to product attributes, product comparison, user experience, or consumer evaluation. Posts were removed when they were duplicate entries, advertisements, bot-generated content, irrelevant news reports, or texts that did not contain product-related discussion. This stage produced a raw social media dataset that was suitable for further text processing.
The second phase transformed raw social media text into clean and structured text data. This step was necessary because social media posts often contain informal language, duplicated content, symbols, links, hashtags, mentions, emojis, and inconsistent spelling. Without preprocessing, these elements may reduce the reliability of the text mining results.
The preprocessing process followed several steps. First, duplicate posts were removed to prevent repeated content from influencing the frequency count. Second, punctuation, symbols, URLs, mentions, hashtags, and irrelevant numbers were removed. Third, all text was converted into lowercase to ensure consistency across terms. Fourth, tokenization was applied to split each post into individual words. Fifth, stopwords were removed because common words do not represent product attributes or consumer evaluations. Sixth, lemmatization and term normalization were conducted to reduce variations of the same meaning.
Term normalization was especially important because the dataset contained product-related words in different forms. For example, “display” and “screen” were treated as visual-related terms, while “battery” and “baterai” were standardized into the same attribute category. After this process, the final output of this phase was a clean text dataset ready for attribute extraction.
The third phase extracted product attributes from the cleaned social media text. A frequency-based text mining approach was used to identify the most dominant terms discussed by consumers. The selected terms were not based only on frequency, but also on their relevance to smartphone product evaluation. This step ensured that the analysis focused on meaningful product attributes rather than general words.
For the iPhone 8, the dominant product-related terms included display, case, iOS, Galaxy, screen, Android, camera, wireless, charger, color, and battery. For the Samsung Galaxy Note 7, the dominant terms included Galaxy, iPhone, charger, color, display, camera, Android, battery, case, wireless, and technology. These terms represent key dimensions of smartphone evaluation, including visual quality, operating system, brand comparison, camera performance, battery performance, charging system, connectivity, design, and technological perception.
After the product attributes were identified, an attribute frequency matrix was constructed. Each row represented an observation unit, while each column represented a product attribute. The value in each cell represented the frequency of a specific attribute in a specific observation unit.
Let denote the frequency of attribute in observation . The attribute frequency matrix can then be expressed as:
where, is the number of observations, is the number of product attributes.
Before PCA was applied, the attribute frequency data were standardized. Standardization was needed because some attributes appeared much more frequently than others. Without standardization, high-frequency terms could dominate the component structure. The z-score transformation was used as follows:
where, is the standardized value of attribute in observation , is the original frequency value, is the mean of attribute , and is the standard deviation of attribute .
The fourth phase applied PCA to reduce the dimensionality of the attribute frequency matrix. PCA was used because product attributes in social media discussions are often correlated. Consumers may discuss display and camera together, battery and charger together, or operating system and ecosystem compatibility together. If these attributes are analyzed separately, the results may contain redundancy and multicollinearity.
PCA transforms correlated attributes into a smaller number of independent principal components. Each component captures a specific pattern of variation in consumer discussions. In this study, the components were interpreted as latent consumer need dimensions because they represent broader structures behind individual product attributes.
The PCA model is expressed as:
where, is the -th principal component, is the loading of attribute on component , and is the standardized value of attribute .
The selection and interpretation of principal components were based on explained variance, cumulative variance, and component loadings. Components with higher explained variance were treated as dominant dimensions of consumer discourse. Component loadings were then used to identify the attributes that contributed most strongly to each component. Attributes with high positive or negative loading values were examined to label each latent component.
PCA was conducted separately for the iPhone 8 and Samsung Galaxy Note 7 datasets. This approach allowed the study to compare how consumer attribute discussions were structured for each product.
The fifth phase examined the relationship between the social media-derived principal components and market success. PCR was used because it allows regression analysis to be conducted using independent component scores rather than correlated original attributes. This method is suitable for the study because it reduces multicollinearity and provides a clearer interpretation of which latent consumer discussion dimensions are associated with market success.
In this study, the principal component scores obtained from PCA were used as independent variables. The dependent variable was the standardized sales indicator. Sales data were standardized to make the product success indicator comparable across observations and products.
The standardized sales indicator was calculated as:
where, is the standardized sales value, is the original sales value, is the mean sales value, and is the standard deviation of sales.
The PCR model is expressed as:
where, is the standardized sales indicator, is the intercept, is the regression coefficient, is the principal component score, and is the error term.
The strength of the relationship between each principal component and market success was evaluated using the correlation coefficient and coefficient of determination. The coefficient of determination was calculated as:
where, is the residual sum of squares and is the total sum of squares. A higher value indicates that a principal component explains a larger proportion of variation in the standardized sales indicator.
The results were interpreted in three stages. First, the dominant terms from text mining were examined to identify the product attributes that received the most consumer attention. These attributes provided an initial view of what consumers discussed when evaluating each smartphone.
Second, the PCA results were interpreted using explained variance and component loadings. Components with high explained variance were considered important because they captured major patterns in consumer discourse. The loading structure was used to understand the meaning of each component. For example, a component with strong loading values on display, camera, and screen may indicate a visual experience dimension, while a component with strong loading values on battery, charger, and wireless may indicate a power and charging dimension.
Third, the PCR results were interpreted by examining the direction and strength of the relationship between each component and standardized sales. Components with high R2 values were considered important because they had stronger associations with market success. Positive coefficients indicated that the component was positively related to sales performance, while negative coefficients indicated an inverse relationship.
Through this interpretation, the framework identifies not only what consumers discuss, but also which consumer discussion dimensions are more closely linked to market success. This makes the framework useful for product development, launch monitoring, competitive benchmarking, and early risk detection.
The proposed framework was evaluated based on its ability to produce meaningful and interpretable product development insights. The evaluation focused on three criteria.
First, the framework should be able to extract relevant product attributes from unstructured social media text.
Second, it should be able to reduce correlated attributes into interpretable latent components through PCA.
Third, it should be able to link these components to a measurable product success indicator through PCR.
The results were also compared with previous social media analytics studies that mainly focused on topic identification, sentiment analysis, complaint mining, or product opportunity discovery. Unlike those approaches, this framework connects consumer attribute discussions with market success. Therefore, the evaluation emphasizes whether the proposed method can extend social media analytics from descriptive insight extraction to empirical product success analysis.
4. Results and Discussion
This section presents the empirical findings of the proposed data-driven social media framework. The discussion follows the research flow: social media data collection and cleansing, text preprocessing, attribute extraction, PCA-based dimensionality reduction, PCR-based market success analysis, and product development interpretation. This order allows the findings to show how unstructured consumer discussions can be transformed into product development insights and then linked to market success.
The first phase focused on collecting social media discussions related to Apple iPhone 8 and Samsung Galaxy Note 7. These products were selected because both represent flagship smartphones from two major competing ecosystems, namely Apple iOS and Samsung Android. Both products also received strong public attention during their launch periods. This condition makes them suitable cases for examining consumer comments, product attribute discussions, and different market outcomes.
The collected posts contained product names, brand-related terms, and smartphone attribute keywords. Posts were retained when they discussed product features, consumer experience, product comparison, or perceived product value. Irrelevant content, duplicate posts, advertisements, bot-like posts, and unrelated news reposts were removed to improve dataset quality.
After the cleansing process, the remaining data represented consumer-generated discussions that were relevant to the selected products. This dataset became the basis for the text mining analysis. This step was important because social media data often contains noise. Without careful filtering, irrelevant posts could distort the identification of consumer-discussed product attributes.
The observation period was defined as six months after the initial market launch of each product. This period was selected because the early launch stage usually generates intensive consumer discussions, product comparisons, usage evaluations, and reactions to product-related issues. For the Apple iPhone 8, tweets were collected from 22 September 2017 to 22 March 2018. For the Samsung Galaxy Note 7, tweets were collected from 19 August 2016 to 19 February 2017. The dataset represents consumer-generated Twitter/X discussions during the early launch period of each product. Because the two products were released in different years, the observation windows were not identical in calendar time, but they were made comparable by using the same six-month post-launch duration for each case.
Product | Platform | Observation Period | Collection Procedure | Initial Collected Tweets | Removed Records After Cleaning | Final Tweets Used in Analysis |
Apple iPhone 8 | Twitter/X | 22 September 2017–22 March 2018 | Keyword-based scraping using “iPhone 8” & “Apple iPhone 8” | 92,438 | 23,764 | 68,674 |
Samsung Galaxy Note 7 | Twitter/X | 19 August 2016–19 February 2017 | Keyword-based scraping using “Samsung Galaxy Note 7” & “Galaxy Note 7” | 128,956 | 35,482 | 93,474 |
Total | Twitter/X | Six months after each product launch | Keyword-based scraping | 221,394 | 59,246 | 162,148 |
The initial scraping process collected 92,438 tweets related to the Apple iPhone 8 and 128,956 tweets related to the Samsung Galaxy Note 7. The higher number of tweets for the Samsung Galaxy Note 7 reflects the strong public attention around the product during its early market period. After data cleaning, 23,764 iPhone 8 records and 35,482 Galaxy Note 7 records were removed. The removed records consisted of duplicate tweets, advertisements, bot-like posts, irrelevant news re-posts, non-English or unreadable texts, and tweets that mentioned the product name without discussing product attributes or consumer evaluation.
The final dataset consisted of 68,674 tweets for the Apple iPhone 8 and 93,474 tweets for the Samsung Galaxy Note 7. Therefore, the total number of tweets used in the analysis was 162,148. This final dataset was used as the basis for text preprocessing, product attribute extraction, construction of the attribute frequency matrix, PCA, and PCR.
Table 2 summarizes the Twitter/X dataset used in this study.
The second phase prepared the raw social media text for analysis. Social media posts often contain informal expressions, spelling variations, hashtags, links, emojis, mentions, and repeated content. These elements can reduce the quality of text mining results if they are not handled properly.
The preprocessing stage removed duplicate data, punctuation, hyperlinks, mentions, hashtags, symbols, and irrelevant numbers. All text was converted into lowercase to ensure consistent term recognition. Tokenization was then applied to split each post into individual terms. Stopwords were removed because they do not represent product attributes or consumer evaluations.
Term normalization was also conducted to reduce variation in words with similar meanings. For example, “display” and “screen” both refer to visual-related attributes, while “battery” and “baterai” refer to the same power-related attribute. This normalization helped build a more consistent attribute dictionary. The output of this phase was a clean text dataset that could be used to identify product attributes and build the attribute frequency matrix.
The third phase extracted product attributes from the cleaned social media data using frequency-based text mining. The analysis identified the terms most frequently discussed by consumers. These terms were then interpreted as product attributes when they were relevant to smartphone evaluation.
Text mining was applied to identify the dominant terms in consumer discussions related to the iPhone 8 and Samsung Galaxy Note 7. The results are presented in Table 3 and Table 4. In this study, high-frequency terms were interpreted as product attributes or product-related themes that received strong attention from consumers on social media. These terms provide an initial view of how consumers evaluated each product and which attributes shaped the focus of online discourse. Text mining is useful for this purpose because it can extract meaningful terms from unstructured user-generated content and map the main themes discussed by consumers (Li et al., 2022).
Term | Count |
Display | 91550 |
Case | 35277 |
IOS | 12515 |
Galaxy | 11224 |
Screen | 8248 |
Android | 7575 |
Camera | 7477 |
Wireless | 7354 |
Charger | 6514 |
Color | 6164 |
Battery | 3818 |
Term | Count |
Galaxy | 237930 |
iPhone | 51871 |
Charger | 8084 |
Color | 6323 |
Display | 99062 |
Camera | 3065 |
Android | 17559 |
Battery | 9179 |
Case | 15184 |
Wireless | 7761 |
Technology | 4930 |
For the iPhone 8, the term display appeared as the most dominant term, with 91,550 occurrences. This result suggests that visual quality was the main focus of consumer discussion. Other frequently mentioned terms, such as case, screen, and camera, show that consumers also paid strong attention to physical design and visual-related features. This finding is consistent with Purnama et al. (2023) that display and screen size are important attributes in smartphone user preferences.
The appearance of iOS also indicates that the operating system played an important role in consumer evaluation. For Apple products, the operating system is not only a technical feature, but also part of the broader product ecosystem. Consumers may evaluate iPhone 8 based on how well the device integrates with Apple services, applications, and accessories.
Functional terms such as wireless, charger, and battery show that consumers also discussed practical usage aspects. These attributes relate to daily use, charging convenience, and device performance. Interestingly, the terms Galaxy and Android also appeared among the dominant terms. This indicates that consumer discussions about the iPhone 8 were comparative. Consumers did not evaluate the product in isolation. They also compared it with Samsung Galaxy products and the Android ecosystem when forming judgments about product value.
For the Samsung Galaxy Note 7, consumer discourse was dominated by Galaxy, with 237,930 occurrences, followed by display, iPhone, and Android. The high frequency of Galaxy shows that brand identity was central in online discussions. This is relevant because brand identity can shape product image, consumer perception, and perceived product differentiation (Dissanayake & Amarasuriya, 2015).
The strong presence of display indicates that consumers associated the Galaxy Note 7 with visual performance and screen quality. This is important because the Galaxy Note series was positioned as a premium smartphone line with a large screen and productivity-oriented design. The term color also suggests that aesthetic appearance contributed to consumer discussion, although it appeared less frequently than technical and brand-related terms.
The relatively high frequency of iPhone and Android confirms that consumers discussed the Galaxy Note 7 in a competitive context. The product was not evaluated only as a Samsung device. It was also compared with Apple products and positioned within the broader Android ecosystem. This pattern shows that consumer evaluation in the smartphone market is strongly shaped by brand rivalry and platform preference.
Functional attributes such as battery, charger, wireless, and camera also appeared in the dominant terms. These terms reflect consumer attention to technical performance and supporting features. The presence of battery is especially important for the Galaxy Note 7 because the product later became associated with battery safety issues. The term technology further indicates that consumers perceived the Galaxy Note 7 as a technology-driven product, not merely as a collection of individual hardware features.
Overall, the text mining results show that consumer discussions for both products were multidimensional. The iPhone 8 discourse was strongly related to display quality, design, ecosystem, and cross-platform comparison. The Samsung Galaxy Note 7 discourse was more strongly centered on brand identity, display quality, Android identity, battery-related issues, and comparison with iPhone. These findings justify the next stage of analysis, where PCA is used to reduce correlated product attributes into latent consumer need components.
PCA was applied to reduce the dimensionality of consumer-discussed product attributes and to identify the main latent dimensions underlying social media discourse. The explained variance results of the first five principal components are presented in Table 5. The proportion of explained variance indicates how much information from the original attribute variables can be represented by each principal component. A higher percentage means that the component captures a larger share of the overall discourse structure.
Component | iPhone8_Var% | Note7_Var% | iPhone8_Cum | Note7_Cum |
PC1 | 49.1 | 65.8 | 49% | 66% |
PC2 | 29.5 | 14.7 | 79% | 81% |
PC3 | 11.5 | 8.7 | 90% | 89% |
PC4 | 5.2 | 6.2 | 95% | 95% |
PC5 | 3.1 | 4.3 | 98% | 100% |
For the iPhone 8, the first principal component explains 49.1% of the total variance. This indicates that almost half of the variation in consumer attribute discussions can be represented by one dominant latent dimension. The second principal component explains an additional 29.5%, increasing the cumulative explained variance to approximately 78.6%, or about 79%. This result shows that the first two components already capture most of the information contained in the original product attribute variables.
The third principal component contributes 11.5% of the variance and increases the cumulative explained variance to approximately 90.1%. This means that the first three components provide a compact but still comprehensive representation of iPhone 8 discourse. The fourth and fifth components contribute smaller additional proportions, 5.2% and 3.1%, respectively. Together, the first five components explain about 98% of the total variance. This pattern suggests that consumer discussions about the iPhone 8 are structured around several main latent dimensions rather than many isolated attributes. These dimensions likely reflect broader consumer concerns, such as visual experience, ecosystem compatibility, design, connectivity, and functional performance.
For the Samsung Galaxy Note 7, the first principal component explains 65.8% of the total variance. This value is substantially higher than the first component of the iPhone 8. It indicates that consumer discourse around the Note 7 was more concentrated in one dominant latent dimension. The second principal component explains 14.7%, bringing the cumulative explained variance to approximately 80.5%, or about 81%. Thus, the first two components already summarize most of the variation in the Note 7 attribute data.
The third component explains 8.7% of the variance, increasing the cumulative proportion to 89%. The fourth and fifth components explain 6.2% and 4.3%, respectively. By the fifth component, the cumulative explained variance reaches almost 100%. This result indicates that the main structure of consumer discourse on the Galaxy Note 7 can be represented by only a few principal components. Compared with the iPhone 8, the Note 7 shows a stronger concentration of variance in the first component. This suggests that consumer discussions about the Note 7 were more strongly shaped by fewer dominant themes.
The PCA results confirm that the original attribute variables can be reduced into a smaller number of meaningful latent components without losing substantial information. For both products, the first two components explain around 80% of the total variance, while the first five components explain nearly all of the variance. This supports the suitability of PCA as a dimensionality reduction method in this study. The resulting component scores can then be used in the next stage of analysis, namely PCR, to examine the relationship between social media-derived latent dimensions and market success
After the explained variance was examined, the next step was to interpret the component loadings. Component loadings indicate the contribution of each attribute to a principal component. Attributes with larger absolute loading values have stronger influence in shaping the meaning of a component. Therefore, this stage helps explain the latent consumer need dimensions behind the principal components.
For the iPhone 8, PC1 is mainly shaped by iOS, USB, charger, Samsung, display, and form (Table 6). The negative loadings on iOS, USB, charger, Samsung, and display indicate that these attributes move in the same direction within this component, while the positive loading on form moves in the opposite direction. This pattern suggests that PC1 captures a latent dimension related to ecosystem compatibility, connectivity, visual quality, and cross-brand comparison.
Feature | PC1 | PC11 |
iOS | -0.418 | +0.233 |
USB | -0.418 | +0.604 |
Charger | -0.363 | -0.480 |
Samsung | -0.342 | -0.269 |
Display | -0.333 | — |
Form | +0.327 | — |
Camera | — | -0.362 |
Speed | — | -0.303 |
The presence of iOS and USB indicates that consumers discussed the iPhone 8 in relation to platform compatibility and device connectivity. Charger also contributes to this component, showing that charging convenience and supporting accessories were part of the same discourse structure. The loading on Samsung shows that consumers evaluated the iPhone 8 in comparison with competing products. Display further strengthens the interpretation that visual experience was an important part of this latent dimension. The opposite loading of form suggests that physical design was discussed as a separate but related consideration within the same component.
PC11 shows a different structure. This component has strong positive loadings on USB and iOS, while charger, camera, speed, and Samsung have negative loadings. The strongest loading is found on USB, followed by charger. This pattern indicates that PC11 reflects a technical utility dimension. More specifically, it contrasts platform and connectivity compatibility with supporting feature performance.
In this component, USB and iOS represent the technical side of compatibility within the Apple ecosystem, while charger, camera, and speed represent practical performance attributes. The negative loading on Samsung again confirms that cross-brand comparison remained part of consumer evaluation. Overall, PC11 captures how consumers linked technical compatibility with functional performance when discussing the iPhone 8.
For the Samsung Galaxy Note 7, PC7 is mainly represented by color, Android, iPhone, camera, wireless, and case (Table 7). The strongest loading is found on color, followed by Android. These two attributes indicate that PC7 reflects aesthetic expression and platform preference. Color represents the visual and design-related side of the product, while Android represents the operating system and ecosystem identity.
Feature | PC7 | PC9 |
Color | +0.634 | — |
Android | +0.414 | — |
iPhone | -0.346 | +0.301 |
Camera | +0.289 | — |
Wireless | -0.288 | -0.207 |
Case | -0.232 | -0.476 |
Technology | — | -0.632 |
Battery | — | +0.355 |
Display | — | +0.210 |
The positive loading on camera shows that imaging capability was also part of this component. This suggests that consumers connected visual appearance, platform identity, and camera performance when discussing the Galaxy Note 7. At the same time, the negative loading on iPhone shows that this component also contains a cross-brand comparison element. Consumers did not discuss the Galaxy Note 7 as an isolated product. They positioned it against the iPhone and the Apple ecosystem. The negative loadings on wireless and case suggest that accessory-related and connectivity-related attributes were evaluated in a different direction from color, Android, and camera.
PC9 shows a more risk-sensitive structure. This component is dominated by technology, case, battery, iPhone, wireless, and display. The strongest loading is the negative loading on technology, followed by the negative loading on case. Battery, iPhone, and display contribute positively to this component. This structure indicates a tension between general technology perception and attributes related to durability, product comparison, and visual performance.
The presence of battery is especially important in the Samsung Galaxy Note 7 case. Unlike ordinary smartphone attributes, battery became a critical issue because of the product’s historical safety problem. The loading pattern suggests that consumers associated technology-related discourse with concerns about supporting attributes and product reliability. The positive loading on iPhone also shows that consumers evaluated the Note 7 against its main competitor when discussing these issues.
Overall, the loading analysis confirms that the principal components represent meaningful consumer discourse dimensions. For the iPhone 8, the key components are related to ecosystem compatibility, connectivity, visual experience, and technical utility. For the Samsung Galaxy Note 7, the key components are related to aesthetic expression, Android platform identity, cross-brand comparison, and risk-related technology perception. These findings provide the basis for the next stage of analysis, where the component scores are linked to standardized sales indicators through PCR. Table 8 summarizes these interpretations across both products.
Product | Principal Component | Dominant Loading Variables | Interpretation Label |
iPhone 8 | PC1 | iOS, USB, Charger, Samsung, Display, Form | Ecosystem compatibility and cross-brand evaluation |
iPhone 8 | PC11 | USB (+), iOS (+), Charger (−), Camera (−), Speed (−) | Technical utility and platform compatibility |
Samsung Galaxy Note 7 | PC7 | Color (+), Android (+), iPhone (−), Camera (+), Wireless (−) | Aesthetic expression and platform identity |
Samsung Galaxy Note 7 | PC9 | Technology (−), Case (−), Battery (+), Display (+) | Risk-sensitive technology and product reliability perception |
After the principal components were obtained and interpreted, the PCR was used to examine the relationship between each component and the standardized success indicator. The component scores were treated as independent variables, while the standardized success indicator was used as the dependent variable. This stage was conducted to identify which latent consumer discussion dimensions were most strongly associated with market success.
The PCR results for the iPhone 8 are presented in Table 9. The table shows the correlation coefficient, coefficient of determination, and regression equation for each principal component.
PC | Correlation | R² | Model |
1 | 0.978 | 0.956 | y = 0.4609x + 1.0675 |
2 | 0.465 | 0.216 | y = 0.5740x + 0.2693 |
3 | 0.309 | 0.096 | y = 0.4483x + 0.3040 |
4 | -0.254 | 0.065 | y = -0.4588x + 0.6481 |
5 | 0.210 | 0.044 | y = 0.4451x + 0.3495 |
6 | 0.366 | 0.134 | y = 2.4509x + 0.9969 |
7 | 0.312 | 0.097 | y = 4.0549x - 0.8270 |
8 | -0.470 | 0.221 | y = -72.598x - 34.251 |
9 | 0.192 | 0.037 | y = 25.259x + 14.741 |
10 | 0.230 | 0.053 | y = 103.58x - 43.35 |
11 | 0.622 | 0.387 | y = 355.36x + 99.366 |
For the iPhone 8, PC1 has the strongest positive association with the standardized success indicator. The correlation value is 0.978, and the R² value is 0.956. This result indicates that PC1 explains 95.6% of the variation in the standardized success indicator. Since PC1 represents a discourse dimension related to ecosystem compatibility, connectivity, display, charging, and cross-brand comparison, the result suggests that these attributes were closely linked to the market performance of the iPhone 8.
This finding is important because it shows that consumer discussions about the iPhone 8 were not only concentrated on individual features. Instead, the strongest market-related dimension combined several connected attributes. Consumers appeared to evaluate the product through the integration of iOS, connectivity, charging support, visual experience, and comparison with Samsung or Android-based products. This pattern is consistent with the positioning of Apple products, where ecosystem integration and user experience often shape consumer perception of product value.
PC11 also shows a meaningful positive relationship with the standardized success indicator. The correlation value is 0.622, and the R² value is 0.387. This means that PC11 explains 38.7% of the variation in the standardized success indicator. Based on the loading structure, PC11 reflects a technical utility dimension that contrasts USB and iOS compatibility with supporting feature performance, such as charger, camera, speed, and Samsung-related comparison. Although its explanatory power is lower than PC1, PC11 still indicates that technical compatibility and functional performance contributed to consumer evaluation of the iPhone 8.
Other components show weaker individual relationships with the standardized success indicator. PC2, PC3, PC5, PC6, PC7, PC9, and PC10 have positive correlations, but their R² values are relatively low. PC4 and PC8 show negative correlations, but their explanatory power is also limited compared with PC1. These results suggest that most individual components outside PC1 and PC11 had weaker roles in explaining iPhone 8 success. Therefore, the market-related discourse for the iPhone 8 was mainly concentrated in one dominant component, supported by one additional technical utility component.
The PCR results for the Samsung Galaxy Note 7 are presented in Table 10. Compared with the iPhone 8, the Note 7 shows a more varied pattern of positive and negative associations across several principal components.
For the Samsung Galaxy Note 7 (Table 10), PC7 has the strongest positive association with the standardized success indicator. The correlation value is 0.943, and the R² value is 0.889. This result indicates that PC7 explains 88.9% of the variation in the standardized success indicator. Based on the loading interpretation, PC7 is associated with color, Android, camera, iPhone, wireless, and case. This component reflects aesthetic expression, Android platform preference, camera-related value, and cross-brand comparison.
The strong positive relationship between PC7 and the success indicator suggests that design-related appeal, platform identity, and camera-related discussion were important positive signals in the Note 7 discourse. Consumers appeared to connect the product with visual differentiation, Android ecosystem preference, and competition with the iPhone. These elements may have supported positive product attention during the early market response.
PC1 also shows a strong positive association with the standardized success indicator, with a correlation value of 0.930 and an R² value of 0.858. This means that PC1 explains 85.8% of the variation in the success indicator. The strength of PC1 indicates that the dominant discourse dimension of the Note 7 was also closely related to market response. This result is consistent with the PCA finding that PC1 captured the largest proportion of variance in Note 7 consumer discussions.
PC | Correlation | R² | Model |
1 | 0.930 | 0.858 | y = 0.8908x - 0.1744 |
2 | -0.643 | 0.413 | y = -0.3793x + 0.6768 |
3 | -0.304 | 0.092 | y = -0.3540x + 0.2113 |
4 | 0.766 | 0.587 | y = 0.7443x + 0.5423 |
5 | -0.138 | 0.019 | y = -0.1773x + 0.2683 |
6 | -0.225 | 0.051 | y = -3.6933x + 2.8782 |
7 | 0.943 | 0.889 | y = 93.933x - 11.54 |
8 | 0.126 | 0.016 | y = 16.848x - 7.4842 |
9 | -0.915 | 0.836 | y = -146.07x - 63.493 |
10 | -0.089 | 0.008 | y = -12.188x - 2.1697 |
11 | -0.442 | 0.195 | y = -75.467x - 2.6074 |
PC4 shows a moderate to strong positive association, with a correlation value of 0.766 and an R² value of 0.587. Although its explanatory power is lower than PC7 and PC1, PC4 still explains more than half of the variation in the standardized success indicator. This indicates that the Note 7 market response was not explained by only one component. Several consumer discourse dimensions contributed to the observed success pattern.
The most important negative result appears in PC9. This component has a correlation value of -0.915 and an R² value of 0.836. The high R² value shows that PC9 has strong explanatory power, but the negative sign indicates an inverse relationship with the standardized success indicator. This means that when the discourse represented by PC9 increased, the success indicator tended to decrease.
This negative direction is meaningful. PC9 includes technology, case, battery, iPhone, wireless, and display-related signals. Among these attributes, battery is especially important because the Samsung Galaxy Note 7 was historically associated with battery safety issues. Therefore, PC9 can be interpreted as a risk-sensitive discourse dimension. It may represent consumer attention to technology-related concerns, product reliability, battery performance, and comparison with competing products. Unlike PC7, which reflects product appeal, PC9 appears to capture product risk or negative market signals.
Other components show weaker relationships. PC2 has a negative correlation of -0.643 and an R² value of 0.413, suggesting a moderate inverse association. PC3, PC5, PC6, PC8, PC10, and PC11 have relatively lower R² values. These results indicate that their individual contribution to explaining the standardized success indicator was limited compared with PC7, PC1, PC4, and PC9.
The PCR results show different patterns between the two products. For the iPhone 8, market success is mainly explained by PC1, which represents ecosystem compatibility, connectivity, display, and cross-brand comparison. This indicates a more concentrated relationship between consumer discourse and product success.
The iPhone 8 case suggests that a single dominant latent dimension was sufficient to explain most of the variation in the standardized success indicator.
For the Samsung Galaxy Note 7, the relationship is more complex. Several components show strong associations with the standardized success indicator. PC7, PC1, and PC4 show positive relationships, while PC9 shows a strong negative relationship. This pattern suggests that consumer discourse around the Note 7 contained both product appeal and product risk signals. Positive discourse was linked to design, platform identity, camera, and brand comparison. Negative discourse was linked to technology and battery-related concerns.
Overall, the PCR results confirm that social media-derived principal components can be used to examine the relationship between consumer attribute discussions and product success. The findings also show that high consumer attention does not always imply positive market value. Some discussion dimensions may signal consumer interest, while others may indicate risk, dissatisfaction, or potential product failure. This distinction is important for product development because firms need to identify not only which attributes are frequently discussed, but also whether those discussions are positively or negatively associated with market outcomes.
These results should be interpreted as associative rather than causal. The regression analysis shows the strength and direction of the relationship between social media-derived components and the standardized success indicator. It does not prove that consumer discussions directly caused product success or failure. However, the results provide useful empirical evidence that online consumer discourse contains measurable signals related to market response.
The findings provide several implications for product development. First, social media discussions can help identify product attributes that receive strong consumer attention during and after product launch. Across the two smartphone cases, consumers repeatedly discussed display quality, camera capability, battery performance, charging system, operating system, wireless features, design, and brand comparison. These attributes represent visible areas of consumer evaluation. They should be monitored closely because they can shape early market response and influence how consumers judge product value.
Second, the PCA results show that consumers do not evaluate product attributes as separate features. Instead, attributes tend to form broader perception dimensions. In the iPhone 8 case, the most influential dimension was related to ecosystem compatibility, connectivity, visual experience, charging support, and competitive comparison. This finding suggests that the market value of the iPhone 8 was strongly associated with the integration of hardware, software, accessories, and the Apple ecosystem. For product managers, this means that product success may depend not only on improving individual features, but also on ensuring that these features work together as a coherent user experience.
Third, the Samsung Galaxy Note 7 case shows that social media discourse can capture both product appeal and product risk. Positive components were associated with design, color, Android platform identity, camera, and comparison with the iPhone. These dimensions reflect the product’s appeal as a premium Android smartphone. However, the negative component was linked to technology and battery-related signals. This pattern is important because the Galaxy Note 7 was historically associated with battery safety issues. The result suggests that social media analytics can help firms detect early warning signals when consumer discussions begin to concentrate around risk-sensitive attributes.
Fourth, the PCR results show that the proposed framework extends beyond descriptive text mining. It does not only identify which attributes consumers discuss most often. It also examines which latent discussion dimensions are associated with market success. This distinction is important for product development. A frequently discussed attribute may reflect interest, dissatisfaction, comparison, or concern. By linking social media-derived components to standardized success indicators, firms can better distinguish between attributes that support market performance and attributes that indicate potential product risk.
In practice, firms can apply the proposed framework at three distinct stages of the product decision process. During product planning, the text mining phase can be used to audit the competitive attribute landscape before a product launch. By identifying the dominant attributes in consumer discourse for competing products, the PCA can inform feature investment priorities and help predict consumer expectations. During launch evaluation, the PCA results can be used to monitor whether consumer discussion is coalescing around positive value dimensions or risk-sensitive attributes. A shift in dominant component scores towards risk-related dimensions can trigger the need for early corrective communication or product intervention. During attribute prioritization, the PCR results can be used to rank which consumer discussion dimensions are most strongly associated with market performance, so that product teams can allocate resources to the attributes and feature clusters that have the highest market relevance. These three application stages make the framework operational beyond academic analysis, and suitable for integration into regular product intelligence workflows.
The findings support the relevance of a data-driven social media framework for product development. The proposed framework shows how unstructured consumer discourse can be transformed into measurable product development insights. Text mining first identified the dominant product attributes discussed by consumers. PCA then reduced overlapping and correlated attributes into interpretable latent components. PCR finally examined the relationship between these components and standardized sales indicators. This sequence provides a systematic analytical path from online consumer discussion to market success analysis.
The results extend prior studies that used social media analytics and online consumer data for product development. Previous research has shown that text mining can identify product attributes, customer opinions, and improvement opportunities from online reviews or social media data. For example, Zhang et al. (2018) used online review mining to support product innovation, while Hong & Wang (2021) applied topic mining and deep learning to summarize customer opinions. Rathore & Ilavarasan (2020) also showed that Twitter/X analytics can capture consumer emotions before and after product launch. These studies confirm that online consumer data is useful for understanding customer response. However, most of them focus on extracting topics, opinions, or sentiments. They do not fully examine how these attributes form latent structures or how those structures relate to market success.
This study addresses that limitation by adding PCA and PCR after the text mining stage. PCA is important because consumer-discussed attributes are often interrelated. For example, smartphone users do not discuss display, camera, battery, charger, operating system, and brand comparison as isolated features. These attributes form broader perception dimensions. In the iPhone 8 case, the dominant latent dimension was related to ecosystem compatibility, connectivity, display, charging support, and cross-brand comparison. This result suggests that the market value of the iPhone 8 was strongly connected to Apple’s integrated ecosystem and user experience. The finding is consistent with the view that product value in technology markets depends not only on individual features, but also on how features work together as a coherent product experience.
The results also contribute to studies on product opportunity mining and complaint-based product improvement. Wang et al. (2023) used social media analytics to mine customer complaints and identify product opportunities. Zhang & Song (2024) combined text mining and decision-making methods to support product improvement in a big data environment. These studies show that consumer-generated text can help firms identify improvement priorities. However, product opportunities and complaints do not always indicate market success. A frequently discussed attribute may reflect satisfaction, dissatisfaction, curiosity, or risk. The PCR results in this study help distinguish these meanings by linking latent components to standardized sales indicators.
The comparison between iPhone 8 and Samsung Galaxy Note 7 shows that product context matters. The iPhone 8 case shows a more positive and concentrated market-related discourse. PC1 produced the strongest relationship with the success indicator, with an R2 value of 0.956. This indicates that the main discourse dimension captured most of the market-related variation. In contrast, the Galaxy Note 7 case showed both positive and negative market-related signals. Some components were positively associated with the success indicator, such as components related to aesthetic appeal, Android identity, camera, and competitive comparison. However, other components showed negative associations, especially those linked to technology and battery-related concerns. This pattern reflects the historical nature of the Galaxy Note 7 case, where strong product appeal existed together with serious product risk.
This finding is important for social media-based product development. High consumer attention does not always represent positive product value. In some cases, high discussion intensity may signal dissatisfaction, safety concerns, or product failure. This is especially relevant for the Galaxy Note 7, where battery-related discourse became a risk-sensitive signal. Therefore, firms should not only monitor the volume of consumer discussions. They also need to examine the structure and direction of those discussions in relation to market outcomes.
The study also extends research on data-driven product planning. Panzner et al. (2024) emphasized the need for structured data analytics pipelines that connect data sources, analytical methods, and product planning decisions. The present framework responds to that need by combining text mining, PCA, and PCR in one analytical process. It does not stop at identifying what consumers discuss. It also identifies which latent consumer perception dimensions have stronger market relevance. This makes the framework more useful for product managers who need to prioritize attributes, evaluate launch response, and detect early product risks.
Overall, the proposed framework can support product development in four ways. First, it helps firms identify product attributes that receive strong consumer attention from social media data. Second, it reduces overlapping attributes into latent consumer perception dimensions, which improves interpretability and reduces redundancy. Third, it links these dimensions to market success indicators, allowing firms to identify which consumer discussions matter most for product performance. Fourth, it supports product risk monitoring by distinguishing positive market signals from negative or risk-sensitive discourse.
These findings show that social media analytics can move beyond descriptive insight extraction. When combined with PCA and PCR, social media data can provide empirical signals for product development and market success analysis. The framework can support product planning, launch monitoring, competitive benchmarking, and early corrective action when consumer discourse begins to concentrate around critical product attributes.
5. Conclusion
This study proposed a data-driven social media framework for product innovation by integrating text mining, PCA, and PCR. The framework was designed to transform unstructured consumer discussions from social media into measurable product innovation insights and to examine their relationship with market success. Using Apple iPhone 8 and Samsung Galaxy Note 7 as empirical cases, this study shows that social media discourse can capture consumer attention toward product attributes, platform comparison, and potential product risks.
The text mining results indicate that consumers discussed several key smartphone attributes across both products, including visual features, operating system, connectivity, battery, charging, camera, platform identity, and cross-brand comparison. These findings confirm that consumer evaluation of smartphones is multidimensional. Consumers do not only assess technical specifications, but also consider ecosystem compatibility, brand identity, daily usability, and product comparison with competing alternatives.
The PCR results further show that selected principal components are strongly associated with standardized sales indicators. For the iPhone 8, PC1 has the strongest relationship with standardized sales, with an R2 value of 0.956. This suggests that the iPhone 8 market response was strongly linked to consumer discussions related to ecosystem compatibility, connectivity, visual experience, charging support, and cross-brand comparison. For the Samsung Galaxy Note 7, PC7 shows the strongest relationship with standardized sales, with an R2value of 0.889. Other components show both positive and negative relationships, indicating that social media discourse can reflect not only product appeal, but also product risk.
The main contribution of this study lies in its integrated analytical framework. Unlike studies that focus only on topic extraction, sentiment analysis, or complaint mining, this study links consumer attribute discussions to market success indicators. This approach allows firms to identify which latent dimensions of consumer discourse have stronger market relevance. It also helps distinguish between attributes that support product value and attributes that may signal consumer concern or product risk.
From a managerial perspective, the proposed framework can support product innovation in several ways. It can help firms identify consumer-prioritized attributes from social media data, reduce overlapping product attributes into interpretable components, monitor product launch responses, compare competing products, and detect early warning signals from consumer discourse. These insights can support evidence-based product development and help firms prioritize innovation efforts based on market-relevant consumer signals.
Despite these contributions, several limitations should be acknowledged. The framework was only tested on two smartphone products. Therefore, its generalizability to other product categories (e.g., consumer electronics, fast-moving consumer goods, and industrial products) needs to be tested, as different product types may have different discourse patterns, which may affect the structure of the components. The study also relied on data from a single collection period and a single social media platform because consumer discourse evolves across the product lifecycle and varies by platform, future research could apply the framework to time-series data and compare results across platforms such as Twitter/X, Reddit, and product review sites. Finally, future validation across non-English language markets and emerging product categories would further strengthen the framework's external validity.
Conceptualization, D.A.P. and R.D.K.; methodology, D.A.P.; software, R.D.K.; validation, D.A.P. and R.D.K.; formal analysis, D.A.P.; investigation, R.D.K.; resources, D.A.P.; data curation, D.A.P.; writing—original draft preparation, D.A.P. and R.D.K.; writing—review and editing, D.A.P. and R.D.K.; visualization, D.A.P.; supervision, R.D.K.; project administration, D.A.P.; funding acquisition, R.D.K. All authors have read and agreed to the published version of the manuscript.
All relevant data are included within the article, and additional data can be provided by the corresponding author upon reasonable request.
The authors declare no conflicts of interest.
