Javascript is required
Cai, M. & Zhang, X. (2025). Consumer preference analysis integrating online reviews: A multiple criteria group approach considering individual stochastic behavior. 4OR, 23(1), 97–122. [Crossref]
Cooper, R. G. (2019). The drivers of success in new-product development. Ind. Mark. Manag., 76, 36–47. [Crossref]
Cui, A. S. & Wu, F. (2016). Utilizing customer knowledge in innovation: Antecedents and impact of customer involvement on new product performance. J. Acad. Mark. Sci., 44(4), 516–538. [Crossref]
Ding, Y., Wu, P., Zhao, J., & Zhou, L. (2025). Forecasting product sales using text mining: A case study in new energy vehicle. Electron. Commer. Res., 25(1), 495–527. [Crossref]
Dissanayake, R. & Amarasuriya, T. (2015). Role of brand identity in developing global brands: A literature-based review on case comparison between apple iPhone vs Samsung smartphone brands. Res. J. Bus. Manag., 2(3), 430–440.
Geissinger, A., Laurell, C., Öberg, C., & Sandström, C. (2023). Social media analytics for innovation management research: A systematic literature review and future research agenda. Technovation, 123, 102712. [Crossref]
Greenacre, M., Groenen, P. J., Hastie, T., d’Enza, A. I., Markos, A., & Tuzhilina, E. (2022). Principal component analysis. Nat. Rev. Methods Primers, 2(1), 100. [Crossref]
Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2019). Multivariate Data Analysis. Cengage Learning.
Han, R., Brennecke, J., Borah, D., & Lam, H. K. S. (2025). The use of social media in different phases of the new product development process: A systematic literature review. R D Manag., 55(1), 108–126. [Crossref]
Harvey, S. & Kou, C.Y. (2013). Collective engagement in creative tasks: The role of evaluation in the creative process in groups. Adm. Sci. Q., 58(3), 346–386. [Crossref]
Hong, M. & Wang, H. (2021). Research on customer opinion summarization using topic mining and deep neural network. Math. Comput. Simul., 185, 88–114. [Crossref]
Klein, M. A. & Şener, F. (2023). Product innovation, diffusion and endogenous growth. Rev. Econ. Dyn., 48, 178–201. [Crossref]
Li, S., Liu, F., Zhang, Y., Zhu, B., Zhu, H., & Yu, Z. (2022). Text mining of user-generated content (UGC) for business applications in e-commerce: A systematic review. Mathematics, 10(19), 3554. [Crossref]
Lin, K. Y. (2018). User experience-based product design for smart production to empower industry 4.0 in the glass recycling circular economy. Comput. Ind. Eng., 125, 729–738. [Crossref]
Morgan, T., Obal, M., & Anokhin, S. (2018). Customer participation and new product performance: Towards the understanding of the mechanisms and key contingencies. Res. Policy, 47(2), 498–510. [Crossref]
Najafi-Tavani, Z., Mousavi, S., Zaefarian, G., & Naudé, P. (2020). Relationship learning and international customer involvement in new product design: The moderating roles of customer dependence and cultural distance. J. Bus. Res., 120, 42–58. [Crossref]
Nasrabadi, M. A., Beauregard, Y., & Ekhlassi, A. (2024). The implication of user-generated content in new product development process: A systematic literature review and future research agenda. Technol. Forecast. Soc. Change, 206, 123551. [Crossref]
Panzner, M., von Enzberg, S., & Dumitrescu, R. (2024). Developing a data analytics toolbox for data-driven product planning: A review and survey methodology. AI EDAM, 38, e18. [Crossref]
Purnama, D. A., Subagyo, & Masruroh, N. A. (2023). Online data-driven concurrent product-process-supply chain design in the early stage of new product development. J. Open Innov. Technol. Mark. Complex., 9(3), 100093. [Crossref]
Rathore, A. K. & Ilavarasan, P. V. (2020). Pre-and post-launch emotions in new product development: Insights from twitter analytics of three products. Int. J. Inf. Manage., 50, 111–127. [Crossref]
Sońta-Drączkowska, E., Cichosz, M., Klimas, P., & Pilewicz, T. (2025). Co-creating innovations with users: A systematic literature review and future research agenda for project management. Eur. Manag. J., 43(2), 321–339. [Crossref]
Tian, Q., Cao, G., & Weerawardena, J. (2024). Strategic use of social media in new product development in B2B firms: The role of absorptive capacity. Ind. Mark. Manag., 120, 132–145. [Crossref]
Ulrich, K. T., Eppinger, S. D., & Yang, M. C. (2008). Product Design and Development. McGraw-Hill Education.
Wang, J., Lai, J. Y., & Lin, Y. H. (2023). Social media analytics for mining customer complaints to explore product opportunities. Comput. Ind. Eng., 178, 109104. [Crossref]
Wu, P., Tang, T., Zhou, L., & Martínez, L. (2024). A decision-support model through online reviews: Consumer preference analysis and product ranking. Inf. Process. Manag., 61(4), 103728. [Crossref]
Xie, X. & Jia, Y. (2016). Consumer involvement in new product development: A case study from the online virtual community. Psychol. Mark., 33(12), 1187–1194. [Crossref]
Yakubu, H. & Kwong, C. K. (2021). Forecasting the importance of product attributes using online customer reviews and Google Trends. Technol. Forecast. Soc. Change, 171, 120983. [Crossref]
Zhang, F. & Song, W. (2024). Product improvement in a big data environment: A novel method based on text mining and large group decision making. Expert Syst. Appl., 245, 123015. [Crossref]
Zhang, H., Rao, H., & Feng, J. (2018). Product innovation based on online review data mining: A case study of Huawei phones. Electron. Commer. Res., 18(1), 3–22. [Crossref]
Search
Open Access
Research article

A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success

Dwi Adi Purnama1,2,
Ratih Dianingtyas Kurnia1,2*
1
Department of Industrial Engineering, Faculty of Industrial Technology, Universitas Islam Indonesia, 55584 Yogyakarta, Indonesia
2
Department of Engineering Management, Faculty of Industrial Technology, Universitas Islam Indonesia, 55584 Yogyakarta, Indonesia
Journal of Research, Innovation and Technologies
|
Volume 5, Issue 2, 2026
|
Pages 223-241
Received: 05-08-2026,
Revised: 06-17-2026,
Accepted: 06-27-2026,
Available online: 06-30-2026
View Full Article|Download PDF

Abstract:

Social media provides a rich source of consumer-generated data that can support product innovation decisions. However, firms still face difficulties in transforming unstructured online discussions into measurable insights that are linked to market success. This study proposes a data-driven social media framework for product innovation by integrating text mining, Principal Component Analysis (PCA), and Principal Component Regression (PCR). The framework identifies dominant consumer attribute discussions, reduces correlated attributes into interpretable latent components, and tests their relationship with standardized sales indicators. The study uses two competitive smartphone products, Apple iPhone 8 and Samsung Galaxy Note 7, as empirical cases. Text mining results show that consumers discussed visual, ecosystem, connectivity, battery, charging, camera, and platform-related attributes across both products. PCA results indicate that a small number of principal components can explain most of the variance in consumer discussions. For the iPhone 8, the first two components explain 79% of the total variance, while for the Galaxy Note 7, they explain 81%. The regression results show strong links between selected social media-derived components and market success. For the iPhone 8, PC1 has the strongest relationship with standardized sales, with an R² of 0.956. For the Galaxy Note 7, PC7 shows the strongest relationship, with an R² of 0.889, while other components show both positive and negative relationships. These findings show that social media discussions can provide early signals of consumer needs, product risks, and market response. The proposed framework offers a systematic tool for attribute prioritization, launch monitoring, and evidence-based product development.
Keywords: Social media analytics, Product development, Text mining, Principal Component Analysis, Principal Component Regression, Consumer attributes, Market success
JEL Classification: C38; D12; M31

1. Introduction

Product innovation is essential for firms that compete in dynamic and technology-driven markets. Rapid changes in customer preferences, shorter product life cycles, and intense competition require firms to develop products that match market expectations. Product success depends not only on technological novelty, but also on the ability of the innovation to diffuse and gain acceptance in the market (Klein & Şener, 2023). Therefore, firms need to understand customer needs accurately and use this understanding as the basis for product development decisions (Lin, 2018; Xie & Jia, 2016).

Despite its importance, product innovation remains risky. Many product ideas fail before they reach commercial success. Purnama et al. (2023) reported that less than one percent of initial product development ideas are commercialized. Cooper (2019) also showed that only a small proportion of product concepts that enter the development and testing process become commercially successful. These failures are often linked to changing market needs, high research and development costs, limited resources, and shorter product life cycles (Najafi-Tavani et al., 2020). These conditions show that firms need more effective methods to identify customer needs at the early stage of new product development (NPD).

Customer need identification plays a central role in NPD. Ulrich et al. (2008) emphasized that a comprehensive understanding of customer needs should begin in the early development stage, including data collection, need interpretation, importance assessment, and opportunity identification. Customer involvement also supports product innovation because it helps firms reduce the risk of developing products that do not fit market expectations (Cui & Wu, 2016; Harvey & Kou, 2013; Morgan et al., 2018). This perspective is consistent with recent studies on user innovation and value co-creation, which view customers as active contributors to idea generation, product design, and innovation development. In NPD, user involvement can help firms identify relevant needs, refine product concepts, and reduce uncertainty during early development stages. However, traditional customer need identification methods, such as surveys, interviews, and focus groups, remain time-consuming and resource-intensive because they depend on structured data collection, participant recruitment, and manual interpretation (Nasrabadi et al., 2024; Sońta-Drączkowska et al., 2025). Recent studies show that user-generated content and social media analytics offer scalable alternatives for capturing customer knowledge from online platforms and transforming it into innovation insights (Geissinger et al., 2023; Nasrabadi et al., 2024). Therefore, firms need faster, more scalable, and more cost-efficient approaches to support customer need identification and data-driven product planning (Panzner et al., 2024; Tian et al., 2024).

Social media offers a promising source of customer insight for product development. Consumers use social media to express expectations, preferences, complaints, and comparisons between competing products. These discussions provide real-time information about product attributes, perceived value, and potential adoption barriers. Recent literature confirms that social media analytics has become an important methodological tool in innovation management because it can capture and analyze user-generated content from online platforms (Geissinger et al., 2023). Social media also supports different phases of NPD, including discovery, development, and launch monitoring (Han et al., 2025). In this context, online consumer discussions can help firms identify emerging needs, detect product issues, and evaluate early market responses.

User-generated content has also received growing attention in NPD research. Nasrabadi et al. (2024) reviewed studies on user-generated content in NPD and identified four major themes, namely its impact on NPD, idea mining, feature derivation, and customer need understanding. This shows that online consumer data is not only useful for marketing analysis, but also for product design and innovation decisions. However, social media data is unstructured, informal, noisy, and difficult to interpret directly. Firms need a systematic method to transform online consumer discourse into measurable and actionable product development insights.

Previous studies have used social media analytics to support product development. Zhang et al. (2018) identified important topics using clustering methods. Hong & Wang (2021) summarized customer opinions to identify product attributes. Other studies applied sentiment analysis to measure satisfaction and dissatisfaction in online discussions (Rathore & Ilavarasan, 2020; Zhang et al., 2018) . Recent studies also show that social media analytics can help firms mine customer complaints and identify product opportunities. Wang et al. (2023) developed a social media analytics method that combines sentiment analysis, topic modeling, topic engagement, and topic emergence to explore product improvement opportunities from customer complaints. Similarly, Zhang & Song (2024) showed that product improvement decisions can be strengthened by combining consumer big data, text mining, and group decision-making methods.

Although these studies show the value of social media for product development, several gaps remain. Many studies focus on topic discovery, sentiment classification, or customer complaint mining. These approaches are useful, but they often remain descriptive. They do not fully explain how product attributes discussed on social media form latent consumer need structures. They also do not sufficiently test how those structures relate to measurable product success. Recent literature on data-driven product planning also stresses the need for transparent analytics pipelines that connect data sources, preprocessing, modeling, and evaluation in a systematic way (Panzner et al., 2024). This gap is important because firms need not only to know what consumers discuss, but also to identify which dimensions of discussion are most related to market outcomes.

This study addresses that gap by proposing a data-driven social media framework for product development. The framework integrates text mining, Principal Component Analysis (PCA), and Principal Component Regression (PCR) to connect consumer attribute discussions with market success. Text mining is used to extract dominant product attributes from social media discussions. PCA reduces correlated attributes into a smaller set of interpretable latent components. The PCR then examines the relationship between these components and a quantitative product success indicator.

This study uses two competitive smartphone products, Apple iPhone 8 and Samsung Galaxy Note 7, as empirical cases. These products provide a relevant context because smartphone markets are highly competitive, attribute-driven, and strongly influenced by consumer discourse. By comparing two competing products, this study shows how different patterns of consumer attribute discussions can be linked to different market outcomes.

This study makes three main contributions. First, it develops a systematic framework that transforms unstructured social media text into measurable product development insights. Second, it applies PCA to identify latent structures among consumer-discussed product attributes, which helps reduce redundancy and multicollinearity in textual attribute data. Third, it links these latent components to product success through regression analysis, allowing firms to identify which dimensions of consumer discourse are most relevant to market performance. The proposed framework can support attribute prioritization, launch monitoring, competitive benchmarking, and early risk detection in product development.

2. Literature Reviews

2.1 Customer Involvement and User-Generated Content in New Product Development

Customer involvement plays an important role in NPD because it helps firms understand market needs, refine product concepts, and reduce uncertainty in product decisions. Morgan et al. (2018) found that customer participation is positively related to NPD performance and that this relationship is mediated by innovativeness. This finding shows that customers do not only act as product users, but also as external knowledge sources that can support product innovation (Morgan et al., 2018).

Recent research also shows that user-generated content has become a relevant data source for NPD. Nasrabadi et al. (2024) reviewed user-generated content studies in NPD and identified several key themes, including idea mining, feature derivation, customer need understanding, and the impact of online user content on product development. These themes show that online consumer data can help firms identify product attributes, detect emerging needs, and support product improvement decisions.

In this context, social media expands the role of customer involvement. It allows firms to observe consumer discussions in real time and at a larger scale than traditional customer research methods. Han et al. (2025) reviewed 110 studies from 2002 to 2023 and showed that social media can support three phases of NPD: discovery, development, and launch. Their study also identified several NPD objectives that can be addressed through social media, including idea generation, customer need identification, product evaluation, and launch monitoring.

2.2 Social Media Analytics for Product Development

Social media analytics has emerged as an important approach in innovation management and product development. It allows firms to capture and analyze user-generated content from online platforms. Geissinger et al. (2023) stated that social media analytics can support customer-focused, market-focused, technology-focused, and society-focused innovation research. Their systematic review also shows that the use of social media analytics in innovation management has grown during the last decade, although the field is still developing.

Social media is useful for product development because consumers often discuss product attributes, usage problems, complaints, comparisons, and expectations on online platforms. These discussions can reveal product strengths and weaknesses that may not be captured through formal surveys or interviews. In business-to-business (B2B) contexts, Tian et al. (2024) showed that strategic use of social media can help firms acquire external knowledge, strengthen absorptive capacity, and improve NPD. This supports the view that social media can function as a knowledge source for product planning and innovation decisions.

However, social media data is difficult to use directly. It is unstructured, informal, noisy, and highly dynamic. This creates a methodological challenge for firms that want to convert online discussions into product development insights. Therefore, a systematic analytics framework is needed to process text data, extract relevant product attributes, reduce data complexity, and connect consumer discourse with measurable product outcomes.

2.3 Text Mining for Product Attribute and Customer Need Identification

Text mining is widely used to extract meaningful information from unstructured consumer data. In product development research, text mining helps identify product attributes, customer opinions, usage problems, and potential improvement opportunities. Zhang et al. (2018) used online review data mining to support product innovation and improvement. Their study shows that online consumer reviews can reveal product-related issues and support product design decisions.

Hong & Wang (2021) developed a customer opinion summarization approach using topic mining and deep neural networks. Their study highlights the importance of extracting product attributes and related opinions from large volumes of customer text. This approach helps firms summarize customer perceptions and identify attribute-level insights that can guide product improvement.

Rathore & Ilavarasan (2020) applied Twitter/X analytics to analyze pre-launch and post-launch emotions in NPD. Their study shows that social media can capture changes in consumer emotion across product launch stages. This is important because consumer reactions before and after launch can indicate acceptance, dissatisfaction, or potential product risk.

Recent studies have moved beyond simple text extraction by combining text mining with other analytical methods. Wang et al. (2023) developed a social media analytics method to mine customer complaints and identify product opportunities. Their method combines preprocessing, sentiment analysis, topic modeling, topic engagement, and opportunity evaluation. This shows that text mining can support product opportunity discovery when it is combined with structured analytical procedures.

Zhang & Song (2024) also proposed a product improvement method in a big data environment by combining text mining and large group decision-making. Their findings show that text mining can improve product improvement decisions when it is integrated with decision-support methods. This confirms that customer text data needs to be transformed into structured knowledge before it can guide product development actions.

2.4 Data-Driven Product Planning and the Need for Dimensionality Reduction

Data-driven product planning requires firms to connect customer data, analytical methods, and product decisions through a clear process. Panzner et al. (2024) argued that successful data-driven product planning requires a structured pipeline that links domain knowledge, data analysis, and product planning tasks. Their study proposes a data analytics toolbox to support product planning, showing the need for systematic methods that can guide firms from raw data to actionable product decisions. This view is also consistent with recent studies on online review-based decision support, which show that customer-generated data can be transformed into attribute-level preference analysis and product ranking models (Cai & Zhang, 2025; Wu et al., 2024).

One major issue in social media-based product analysis is the presence of many correlated product attributes. Consumers often discuss attributes together, such as display and camera, battery and charging, or operating system and ecosystem compatibility. If these variables are analyzed separately, the model may suffer from redundancy and multicollinearity. Multicollinearity can reduce model stability, inflate estimation uncertainty, and weaken the interpretability of individual predictors (Greenacre et al., 2022; Hair et al., 2019). In product development research, this issue is important because consumer needs are often multidimensional and interrelated, so single-attribute analysis may fail to capture the broader structure of customer perception (Wu et al., 2024; Zhang & Song, 2024).

The PCA can address this issue by reducing correlated attributes into a smaller set of independent components. PCA transforms a cases-by-variables data table into a smaller number of principal components that retain the main information in the original variables while improving interpretability (Greenacre et al., 2022). Each principal component represents a latent structure in consumer discussions. In product development research, this is useful because consumer needs are often multidimensional and cannot be fully represented by single attributes. PCA helps identify broader dimensions of customer perception, such as visual experience, ecosystem compatibility, charging performance, design preference, or platform identity. Therefore, PCA provides a suitable dimensionality reduction method for converting complex social media-based product attributes into interpretable latent components before further regression analysis.

2.5 Linking Social Media-Derived Attributes to Market Success

Many previous studies use social media analytics to identify product topics, sentiments, complaints, or customer needs. These studies provide valuable descriptive insights for product development, especially in identifying product opportunities, customer complaints, and improvement priorities (Wang et al., 2023; Wu et al., 2024). However, fewer studies examine how social media-derived attributes relate to market success.

This creates an important research gap. Firms need to know not only what consumers discuss, but also which dimensions of discussion are associated with product performance. Prior studies on online reviews show that consumer-generated text can support product sales forecasting and attribute importance prediction, indicating that online discourse can contain signals related to market response (Ding et al., 2025; Yakubu & Kwong, 2021).

The PCR can help address this gap. PCR combines PCA and regression analysis by first transforming correlated predictors into principal components and then using selected components as predictors in a regression model (Greenacre et al., 2022). This approach is suitable for social media-based product development because consumer-discussed attributes are often correlated. PCA reduces these correlated attributes into independent components, while regression tests the relationship between these components and a product success indicator. Therefore, PCR can reduce multicollinearity and help identify which latent dimensions of consumer discourse have stronger links to market outcomes.

In this study, market success is represented through a standardized sales indicator. This allows the study to examine whether consumer attribute discussions on social media can explain variation in product success. By applying PCR, the framework moves beyond descriptive text mining and provides a predictive link between online discourse and product performance. This position is consistent with recent data-driven product development studies, which emphasize the need to transform customer-generated data into measurable decision-support variables for product improvement and market-oriented decision-making (Panzner et al., 2024; Zhang & Song, 2024)

2.6 Research Gap and Position of This Study

Based on the previous studies and Table 1, social media analytics has been widely used to support product development, especially for identifying product topics, customer needs, complaints, emotions, and improvement opportunities. Most studies apply text mining, topic modeling, sentiment analysis, or customer complaint mining to extract insights from online consumer discussions. These approaches are useful because they help firms understand what consumers discuss and how consumers respond to product attributes.

Table 1. Positioning of previous studies and research gap

Author(s)

Study Case

Data Type

Data Source

Text
Analytics

Advanced Modeling

Research Focus

Zhang et al. (2018)

Huawei

UGC

OR

TM

PD

Rathore & Ilavarasan (2020)

3 products

UGC

SM

TM + SA

MS + PD

Hong & Wan (2021)

Product reviews

UGC

OR

TM + Topic

PD

Yakubu & Kwong (2021)

Reviews + Trends

UGC

OR

TM

Reg

PD

Geissinger et al. (2023)

Innovation mgmt SLR

SLR + UGC

SM

PD

Wang et al. (2023)

Complaints mining

UGC

SM

TM + Topic + SA

PD

Panzner et al. (2024)

Data toolbox SLR

SLR + UGC

PD

Nasrabadi et al. (2024)

UGC in NPD SLR

SLR + UGC

SM+OR

PD

Tian et al. (2024)

B2B SM use

UGC

SM

PD

Zhang & Song (2024)

Big data improv.

UGC

OR

TM

PD

Han et al. (2025)

SM across NPD SLR

SLR + UGC

SM

PD

Ding et al. (2025)

NEV sales forecast

UGC

OR

TM + SA

Reg

MS + PD

This study

iPhone 8 & Note7

UGC

SM

TM

PCA + Reg

MS + PD

TM = Text Mining; Topic = Topic Modeling; SA = Sentiment Analysis; PCA = Principal Component Analysis; Reg = Regression;

SM = Social Media; OR = Online Review; MS = Market Success; PD = Product Development; SLR = Systematic Literature Review;

UGC = user-generated content; en dash “–” indicates the corresponding method or focus was not covered in the cited study.

However, three important gaps remain. First, many studies still focus on descriptive analysis. They identify topics, sentiments, or complaints, but do not explain the latent structure among product attributes discussed by consumers. In social media discourse, product attributes are often correlated. For example, consumers may discuss display, camera, battery, operating system, and charging features together. If these attributes are analyzed separately, the results may contain redundancy and multicollinearity.

Second, prior studies rarely connect social media-derived product attributes with measurable market success. Most existing studies stop at identifying customer opinions or product improvement opportunities. They do not test whether the extracted attributes are related to product performance indicators, such as sales or market response. This limits the practical value of social media analytics for product development decisions.

Third, there is still limited integration between text mining, dimensionality reduction, and regression analysis in one systematic product development framework. Existing approaches often use text mining or sentiment analysis as standalone tools. A more integrated method is needed to transform unstructured social media text into latent consumer need components and then examine their relationship with product success.

Therefore, this study addresses these gaps by proposing a data-driven social media framework for product development that integrates text mining, PCA and PCR. The framework identifies consumer attribute discussions, reduces correlated attributes into interpretable latent components, and links these components to market success. The proposed framework differs from previous studies in social media analytics in three main aspects. Methodologically, it combines text mining, PCA and PCR into one single pipeline, rather than applying each method as a separate tool. Previous studies have typically used text mining or topic modelling to identify product attributes or consumer sentiments without examining how those attributes form latent structures or how those structures relate to quantifiable market outcomes. The research extends beyond providing descriptive insights at the application level. Instead of only identifying what consumers discuss, the framework examines which latent discussion dimensions are most strongly associated with market success, distinguishing between attributes that signal product appeal and those that indicate product risk. This in turn makes the framework suitable for attribute discovery, launch monitoring, competitive benchmarking and early risk detection in product development.

3. Methodology

This study develops a data-driven social media framework for product development by linking consumer attribute discussions with market success. The framework uses social media text as the main empirical input because online consumer discussions contain direct expressions of product expectations, complaints, comparisons, and perceived product attributes. The method combines text mining, PCA and PCR to transform unstructured consumer discourse into measurable product development insights.

The proposed framework consists of five main phases (Figure 1). The first phase collects and cleans social media data related to the selected products. The second phase preprocesses the text data to prepare it for analysis. The third phase extracts product attributes and builds an attribute frequency matrix. The fourth phase applies PCA to reduce correlated attributes into latent consumer need components. The fifth phase uses PCR to examine the relationship between these components and market success. The framework is applied to two competitive smartphone products, Apple iPhone 8 and Samsung Galaxy Note 7.

Figure 1. Overview of the proposed data-driven social media framework for product development
3.1 Phase 1: Online Data Mining and Cleansing

The first phase collected consumer discussions from social media related to Apple iPhone 8 and Samsung Galaxy Note 7. These two products were selected based on three historical considerations. First, both products represent flagship smartphones from two major competing ecosystems, Apple iOS and Samsung Android, which often generate intense consumer comparisons in terms of display, camera, battery, operating system, charging, design, and brand preference. Second, both products had strong public visibility during their launch periods, making them suitable for identifying consumer comments from online discussions. Third, the two products provide contrasting historical market outcomes. Therefore, the comparison between iPhone 8 and Samsung Galaxy Note 7 provides a relevant empirical setting to examine how different patterns of consumer attribute discussions relate to different market success outcomes.

Data collection focused on posts that contained product names, competing brand terms, and smartphone-related attributes. The search keywords included product-related terms such as “iPhone 8”, “Apple iPhone 8”, “Samsung Galaxy Note 7”, “Galaxy Note 7”.

To improve data quality, several cleaning criteria were applied. Posts were retained when they mentioned the selected products and contained discussions related to product attributes, product comparison, user experience, or consumer evaluation. Posts were removed when they were duplicate entries, advertisements, bot-generated content, irrelevant news reports, or texts that did not contain product-related discussion. This stage produced a raw social media dataset that was suitable for further text processing.

3.2 Phase 2: Text Data Preprocessing

The second phase transformed raw social media text into clean and structured text data. This step was necessary because social media posts often contain informal language, duplicated content, symbols, links, hashtags, mentions, emojis, and inconsistent spelling. Without preprocessing, these elements may reduce the reliability of the text mining results.

The preprocessing process followed several steps. First, duplicate posts were removed to prevent repeated content from influencing the frequency count. Second, punctuation, symbols, URLs, mentions, hashtags, and irrelevant numbers were removed. Third, all text was converted into lowercase to ensure consistency across terms. Fourth, tokenization was applied to split each post into individual words. Fifth, stopwords were removed because common words do not represent product attributes or consumer evaluations. Sixth, lemmatization and term normalization were conducted to reduce variations of the same meaning.

Term normalization was especially important because the dataset contained product-related words in different forms. For example, “display” and “screen” were treated as visual-related terms, while “battery” and “baterai” were standardized into the same attribute category. After this process, the final output of this phase was a clean text dataset ready for attribute extraction.

3.3 Phase 3: Product Attribute Extraction and Frequency Matrix Construction

The third phase extracted product attributes from the cleaned social media text. A frequency-based text mining approach was used to identify the most dominant terms discussed by consumers. The selected terms were not based only on frequency, but also on their relevance to smartphone product evaluation. This step ensured that the analysis focused on meaningful product attributes rather than general words.

For the iPhone 8, the dominant product-related terms included display, case, iOS, Galaxy, screen, Android, camera, wireless, charger, color, and battery. For the Samsung Galaxy Note 7, the dominant terms included Galaxy, iPhone, charger, color, display, camera, Android, battery, case, wireless, and technology. These terms represent key dimensions of smartphone evaluation, including visual quality, operating system, brand comparison, camera performance, battery performance, charging system, connectivity, design, and technological perception.

After the product attributes were identified, an attribute frequency matrix was constructed. Each row represented an observation unit, while each column represented a product attribute. The value in each cell represented the frequency of a specific attribute in a specific observation unit.

Let denote the frequency of attribute in observation . The attribute frequency matrix can then be expressed as:

where, is the number of observations, is the number of product attributes.

Before PCA was applied, the attribute frequency data were standardized. Standardization was needed because some attributes appeared much more frequently than others. Without standardization, high-frequency terms could dominate the component structure. The z-score transformation was used as follows:

where, is the standardized value of attribute in observation , is the original frequency value, is the mean of attribute , and is the standard deviation of attribute .

3.4 Phase 4: Latent Component Modeling Using Principal Component Analysis

The fourth phase applied PCA to reduce the dimensionality of the attribute frequency matrix. PCA was used because product attributes in social media discussions are often correlated. Consumers may discuss display and camera together, battery and charger together, or operating system and ecosystem compatibility together. If these attributes are analyzed separately, the results may contain redundancy and multicollinearity.

PCA transforms correlated attributes into a smaller number of independent principal components. Each component captures a specific pattern of variation in consumer discussions. In this study, the components were interpreted as latent consumer need dimensions because they represent broader structures behind individual product attributes.

The PCA model is expressed as:

where, is the -th principal component, is the loading of attribute on component , and is the standardized value of attribute .

The selection and interpretation of principal components were based on explained variance, cumulative variance, and component loadings. Components with higher explained variance were treated as dominant dimensions of consumer discourse. Component loadings were then used to identify the attributes that contributed most strongly to each component. Attributes with high positive or negative loading values were examined to label each latent component.

PCA was conducted separately for the iPhone 8 and Samsung Galaxy Note 7 datasets. This approach allowed the study to compare how consumer attribute discussions were structured for each product.

3.5 Phase 5: Market Success Analysis Using Principal Component Regression

The fifth phase examined the relationship between the social media-derived principal components and market success. PCR was used because it allows regression analysis to be conducted using independent component scores rather than correlated original attributes. This method is suitable for the study because it reduces multicollinearity and provides a clearer interpretation of which latent consumer discussion dimensions are associated with market success.

In this study, the principal component scores obtained from PCA were used as independent variables. The dependent variable was the standardized sales indicator. Sales data were standardized to make the product success indicator comparable across observations and products.

The standardized sales indicator was calculated as:

where, is the standardized sales value, is the original sales value, is the mean sales value, and is the standard deviation of sales.

The PCR model is expressed as:

where, is the standardized sales indicator, is the intercept, is the regression coefficient, is the principal component score, and is the error term.

The strength of the relationship between each principal component and market success was evaluated using the correlation coefficient and coefficient of determination. The coefficient of determination was calculated as:

where, is the residual sum of squares and is the total sum of squares. A higher value indicates that a principal component explains a larger proportion of variation in the standardized sales indicator.

3.6 Interpretation of Product Development Insights

The results were interpreted in three stages. First, the dominant terms from text mining were examined to identify the product attributes that received the most consumer attention. These attributes provided an initial view of what consumers discussed when evaluating each smartphone.

Second, the PCA results were interpreted using explained variance and component loadings. Components with high explained variance were considered important because they captured major patterns in consumer discourse. The loading structure was used to understand the meaning of each component. For example, a component with strong loading values on display, camera, and screen may indicate a visual experience dimension, while a component with strong loading values on battery, charger, and wireless may indicate a power and charging dimension.

Third, the PCR results were interpreted by examining the direction and strength of the relationship between each component and standardized sales. Components with high R2 values were considered important because they had stronger associations with market success. Positive coefficients indicated that the component was positively related to sales performance, while negative coefficients indicated an inverse relationship.

Through this interpretation, the framework identifies not only what consumers discuss, but also which consumer discussion dimensions are more closely linked to market success. This makes the framework useful for product development, launch monitoring, competitive benchmarking, and early risk detection.

3.7 Evaluation of the Proposed Framework

The proposed framework was evaluated based on its ability to produce meaningful and interpretable product development insights. The evaluation focused on three criteria.

First, the framework should be able to extract relevant product attributes from unstructured social media text.

Second, it should be able to reduce correlated attributes into interpretable latent components through PCA.

Third, it should be able to link these components to a measurable product success indicator through PCR.

The results were also compared with previous social media analytics studies that mainly focused on topic identification, sentiment analysis, complaint mining, or product opportunity discovery. Unlike those approaches, this framework connects consumer attribute discussions with market success. Therefore, the evaluation emphasizes whether the proposed method can extend social media analytics from descriptive insight extraction to empirical product success analysis.

4. Results and Discussion

This section presents the empirical findings of the proposed data-driven social media framework. The discussion follows the research flow: social media data collection and cleansing, text preprocessing, attribute extraction, PCA-based dimensionality reduction, PCR-based market success analysis, and product development interpretation. This order allows the findings to show how unstructured consumer discussions can be transformed into product development insights and then linked to market success.

4.1 Social Media Data Collection and Cleansing

The first phase focused on collecting social media discussions related to Apple iPhone 8 and Samsung Galaxy Note 7. These products were selected because both represent flagship smartphones from two major competing ecosystems, namely Apple iOS and Samsung Android. Both products also received strong public attention during their launch periods. This condition makes them suitable cases for examining consumer comments, product attribute discussions, and different market outcomes.

The collected posts contained product names, brand-related terms, and smartphone attribute keywords. Posts were retained when they discussed product features, consumer experience, product comparison, or perceived product value. Irrelevant content, duplicate posts, advertisements, bot-like posts, and unrelated news reposts were removed to improve dataset quality.

After the cleansing process, the remaining data represented consumer-generated discussions that were relevant to the selected products. This dataset became the basis for the text mining analysis. This step was important because social media data often contains noise. Without careful filtering, irrelevant posts could distort the identification of consumer-discussed product attributes.

The observation period was defined as six months after the initial market launch of each product. This period was selected because the early launch stage usually generates intensive consumer discussions, product comparisons, usage evaluations, and reactions to product-related issues. For the Apple iPhone 8, tweets were collected from 22 September 2017 to 22 March 2018. For the Samsung Galaxy Note 7, tweets were collected from 19 August 2016 to 19 February 2017. The dataset represents consumer-generated Twitter/X discussions during the early launch period of each product. Because the two products were released in different years, the observation windows were not identical in calendar time, but they were made comparable by using the same six-month post-launch duration for each case.

Table 2. Summary of Twitter/X dataset

Product

Platform

Observation Period

Collection Procedure

Initial Collected Tweets

Removed Records After Cleaning

Final Tweets Used in Analysis

Apple iPhone 8

Twitter/X

22 September 2017–22 March 2018

Keyword-based scraping using “iPhone 8” & “Apple iPhone 8”

92,438

23,764

68,674

Samsung Galaxy Note 7

Twitter/X

19 August 2016–19 February 2017

Keyword-based scraping using “Samsung Galaxy Note 7” & “Galaxy Note 7”

128,956

35,482

93,474

Total

Twitter/X

Six months after each product launch

Keyword-based scraping

221,394

59,246

162,148

The initial scraping process collected 92,438 tweets related to the Apple iPhone 8 and 128,956 tweets related to the Samsung Galaxy Note 7. The higher number of tweets for the Samsung Galaxy Note 7 reflects the strong public attention around the product during its early market period. After data cleaning, 23,764 iPhone 8 records and 35,482 Galaxy Note 7 records were removed. The removed records consisted of duplicate tweets, advertisements, bot-like posts, irrelevant news re-posts, non-English or unreadable texts, and tweets that mentioned the product name without discussing product attributes or consumer evaluation.

The final dataset consisted of 68,674 tweets for the Apple iPhone 8 and 93,474 tweets for the Samsung Galaxy Note 7. Therefore, the total number of tweets used in the analysis was 162,148. This final dataset was used as the basis for text preprocessing, product attribute extraction, construction of the attribute frequency matrix, PCA, and PCR.

Table 2 summarizes the Twitter/X dataset used in this study.

4.2 Text Data Preprocessing

The second phase prepared the raw social media text for analysis. Social media posts often contain informal expressions, spelling variations, hashtags, links, emojis, mentions, and repeated content. These elements can reduce the quality of text mining results if they are not handled properly.

The preprocessing stage removed duplicate data, punctuation, hyperlinks, mentions, hashtags, symbols, and irrelevant numbers. All text was converted into lowercase to ensure consistent term recognition. Tokenization was then applied to split each post into individual terms. Stopwords were removed because they do not represent product attributes or consumer evaluations.

Term normalization was also conducted to reduce variation in words with similar meanings. For example, “display” and “screen” both refer to visual-related attributes, while “battery” and “baterai” refer to the same power-related attribute. This normalization helped build a more consistent attribute dictionary. The output of this phase was a clean text dataset that could be used to identify product attributes and build the attribute frequency matrix.

4.3 Product Attribute Extraction

The third phase extracted product attributes from the cleaned social media data using frequency-based text mining. The analysis identified the terms most frequently discussed by consumers. These terms were then interpreted as product attributes when they were relevant to smartphone evaluation.

Text mining was applied to identify the dominant terms in consumer discussions related to the iPhone 8 and Samsung Galaxy Note 7. The results are presented in Table 3 and Table 4. In this study, high-frequency terms were interpreted as product attributes or product-related themes that received strong attention from consumers on social media. These terms provide an initial view of how consumers evaluated each product and which attributes shaped the focus of online discourse. Text mining is useful for this purpose because it can extract meaningful terms from unstructured user-generated content and map the main themes discussed by consumers (Li et al., 2022).

Table 3. Top terms from text mining for iPhone 8

Term

Count

Display

91550

Case

35277

IOS

12515

Galaxy

11224

Screen

8248

Android

7575

Camera

7477

Wireless

7354

Charger

6514

Color

6164

Battery

3818

Table 4. Top terms from text mining for Samsung Galaxy Note 7

Term

Count

Galaxy

237930

iPhone

51871

Charger

8084

Color

6323

Display

99062

Camera

3065

Android

17559

Battery

9179

Case

15184

Wireless

7761

Technology

4930

For the iPhone 8, the term display appeared as the most dominant term, with 91,550 occurrences. This result suggests that visual quality was the main focus of consumer discussion. Other frequently mentioned terms, such as case, screen, and camera, show that consumers also paid strong attention to physical design and visual-related features. This finding is consistent with Purnama et al. (2023) that display and screen size are important attributes in smartphone user preferences.

The appearance of iOS also indicates that the operating system played an important role in consumer evaluation. For Apple products, the operating system is not only a technical feature, but also part of the broader product ecosystem. Consumers may evaluate iPhone 8 based on how well the device integrates with Apple services, applications, and accessories.

Functional terms such as wireless, charger, and battery show that consumers also discussed practical usage aspects. These attributes relate to daily use, charging convenience, and device performance. Interestingly, the terms Galaxy and Android also appeared among the dominant terms. This indicates that consumer discussions about the iPhone 8 were comparative. Consumers did not evaluate the product in isolation. They also compared it with Samsung Galaxy products and the Android ecosystem when forming judgments about product value.

For the Samsung Galaxy Note 7, consumer discourse was dominated by Galaxy, with 237,930 occurrences, followed by display, iPhone, and Android. The high frequency of Galaxy shows that brand identity was central in online discussions. This is relevant because brand identity can shape product image, consumer perception, and perceived product differentiation (Dissanayake & Amarasuriya, 2015).

The strong presence of display indicates that consumers associated the Galaxy Note 7 with visual performance and screen quality. This is important because the Galaxy Note series was positioned as a premium smartphone line with a large screen and productivity-oriented design. The term color also suggests that aesthetic appearance contributed to consumer discussion, although it appeared less frequently than technical and brand-related terms.

The relatively high frequency of iPhone and Android confirms that consumers discussed the Galaxy Note 7 in a competitive context. The product was not evaluated only as a Samsung device. It was also compared with Apple products and positioned within the broader Android ecosystem. This pattern shows that consumer evaluation in the smartphone market is strongly shaped by brand rivalry and platform preference.

Functional attributes such as battery, charger, wireless, and camera also appeared in the dominant terms. These terms reflect consumer attention to technical performance and supporting features. The presence of battery is especially important for the Galaxy Note 7 because the product later became associated with battery safety issues. The term technology further indicates that consumers perceived the Galaxy Note 7 as a technology-driven product, not merely as a collection of individual hardware features.

Overall, the text mining results show that consumer discussions for both products were multidimensional. The iPhone 8 discourse was strongly related to display quality, design, ecosystem, and cross-platform comparison. The Samsung Galaxy Note 7 discourse was more strongly centered on brand identity, display quality, Android identity, battery-related issues, and comparison with iPhone. These findings justify the next stage of analysis, where PCA is used to reduce correlated product attributes into latent consumer need components.

4.4 Principal Component Analysis Results

PCA was applied to reduce the dimensionality of consumer-discussed product attributes and to identify the main latent dimensions underlying social media discourse. The explained variance results of the first five principal components are presented in Table 5. The proportion of explained variance indicates how much information from the original attribute variables can be represented by each principal component. A higher percentage means that the component captures a larger share of the overall discourse structure.

Table 5. Explained variance of the first five principal components

Component

iPhone8_Var%

Note7_Var%

iPhone8_Cum

Note7_Cum

PC1

49.1

65.8

49%

66%

PC2

29.5

14.7

79%

81%

PC3

11.5

8.7

90%

89%

PC4

5.2

6.2

95%

95%

PC5

3.1

4.3

98%

100%

PC = Principal Component.

For the iPhone 8, the first principal component explains 49.1% of the total variance. This indicates that almost half of the variation in consumer attribute discussions can be represented by one dominant latent dimension. The second principal component explains an additional 29.5%, increasing the cumulative explained variance to approximately 78.6%, or about 79%. This result shows that the first two components already capture most of the information contained in the original product attribute variables.

The third principal component contributes 11.5% of the variance and increases the cumulative explained variance to approximately 90.1%. This means that the first three components provide a compact but still comprehensive representation of iPhone 8 discourse. The fourth and fifth components contribute smaller additional proportions, 5.2% and 3.1%, respectively. Together, the first five components explain about 98% of the total variance. This pattern suggests that consumer discussions about the iPhone 8 are structured around several main latent dimensions rather than many isolated attributes. These dimensions likely reflect broader consumer concerns, such as visual experience, ecosystem compatibility, design, connectivity, and functional performance.

For the Samsung Galaxy Note 7, the first principal component explains 65.8% of the total variance. This value is substantially higher than the first component of the iPhone 8. It indicates that consumer discourse around the Note 7 was more concentrated in one dominant latent dimension. The second principal component explains 14.7%, bringing the cumulative explained variance to approximately 80.5%, or about 81%. Thus, the first two components already summarize most of the variation in the Note 7 attribute data.

The third component explains 8.7% of the variance, increasing the cumulative proportion to 89%. The fourth and fifth components explain 6.2% and 4.3%, respectively. By the fifth component, the cumulative explained variance reaches almost 100%. This result indicates that the main structure of consumer discourse on the Galaxy Note 7 can be represented by only a few principal components. Compared with the iPhone 8, the Note 7 shows a stronger concentration of variance in the first component. This suggests that consumer discussions about the Note 7 were more strongly shaped by fewer dominant themes.

The PCA results confirm that the original attribute variables can be reduced into a smaller number of meaningful latent components without losing substantial information. For both products, the first two components explain around 80% of the total variance, while the first five components explain nearly all of the variance. This supports the suitability of PCA as a dimensionality reduction method in this study. The resulting component scores can then be used in the next stage of analysis, namely PCR, to examine the relationship between social media-derived latent dimensions and market success

4.5 Interpretation of Component Loadings

After the explained variance was examined, the next step was to interpret the component loadings. Component loadings indicate the contribution of each attribute to a principal component. Attributes with larger absolute loading values have stronger influence in shaping the meaning of a component. Therefore, this stage helps explain the latent consumer need dimensions behind the principal components.

For the iPhone 8, PC1 is mainly shaped by iOS, USB, charger, Samsung, display, and form (Table 6). The negative loadings on iOS, USB, charger, Samsung, and display indicate that these attributes move in the same direction within this component, while the positive loading on form moves in the opposite direction. This pattern suggests that PC1 captures a latent dimension related to ecosystem compatibility, connectivity, visual quality, and cross-brand comparison.

Table 6. Leading component loadings for iPhone 8

Feature

PC1

PC11

iOS

-0.418

+0.233

USB

-0.418

+0.604

Charger

-0.363

-0.480

Samsung

-0.342

-0.269

Display

-0.333

Form

+0.327

Camera

-0.362

Speed

-0.303

PC = Principal Component; Em dash (—) denotes loadings not reported (absolute values below the threshold).

The presence of iOS and USB indicates that consumers discussed the iPhone 8 in relation to platform compatibility and device connectivity. Charger also contributes to this component, showing that charging convenience and supporting accessories were part of the same discourse structure. The loading on Samsung shows that consumers evaluated the iPhone 8 in comparison with competing products. Display further strengthens the interpretation that visual experience was an important part of this latent dimension. The opposite loading of form suggests that physical design was discussed as a separate but related consideration within the same component.

PC11 shows a different structure. This component has strong positive loadings on USB and iOS, while charger, camera, speed, and Samsung have negative loadings. The strongest loading is found on USB, followed by charger. This pattern indicates that PC11 reflects a technical utility dimension. More specifically, it contrasts platform and connectivity compatibility with supporting feature performance.

In this component, USB and iOS represent the technical side of compatibility within the Apple ecosystem, while charger, camera, and speed represent practical performance attributes. The negative loading on Samsung again confirms that cross-brand comparison remained part of consumer evaluation. Overall, PC11 captures how consumers linked technical compatibility with functional performance when discussing the iPhone 8.

For the Samsung Galaxy Note 7, PC7 is mainly represented by color, Android, iPhone, camera, wireless, and case (Table 7). The strongest loading is found on color, followed by Android. These two attributes indicate that PC7 reflects aesthetic expression and platform preference. Color represents the visual and design-related side of the product, while Android represents the operating system and ecosystem identity.

Table 7. Leading component loadings for Samsung Galaxy Note 7

Feature

PC7

PC9

Color

+0.634

Android

+0.414

iPhone

-0.346

+0.301

Camera

+0.289

Wireless

-0.288

-0.207

Case

-0.232

-0.476

Technology

-0.632

Battery

+0.355

Display

+0.210

PC = Principal Component; Em dash (—) denotes loadings not reported (absolute values below the threshold).

The positive loading on camera shows that imaging capability was also part of this component. This suggests that consumers connected visual appearance, platform identity, and camera performance when discussing the Galaxy Note 7. At the same time, the negative loading on iPhone shows that this component also contains a cross-brand comparison element. Consumers did not discuss the Galaxy Note 7 as an isolated product. They positioned it against the iPhone and the Apple ecosystem. The negative loadings on wireless and case suggest that accessory-related and connectivity-related attributes were evaluated in a different direction from color, Android, and camera.

PC9 shows a more risk-sensitive structure. This component is dominated by technology, case, battery, iPhone, wireless, and display. The strongest loading is the negative loading on technology, followed by the negative loading on case. Battery, iPhone, and display contribute positively to this component. This structure indicates a tension between general technology perception and attributes related to durability, product comparison, and visual performance.

The presence of battery is especially important in the Samsung Galaxy Note 7 case. Unlike ordinary smartphone attributes, battery became a critical issue because of the product’s historical safety problem. The loading pattern suggests that consumers associated technology-related discourse with concerns about supporting attributes and product reliability. The positive loading on iPhone also shows that consumers evaluated the Note 7 against its main competitor when discussing these issues.

Overall, the loading analysis confirms that the principal components represent meaningful consumer discourse dimensions. For the iPhone 8, the key components are related to ecosystem compatibility, connectivity, visual experience, and technical utility. For the Samsung Galaxy Note 7, the key components are related to aesthetic expression, Android platform identity, cross-brand comparison, and risk-related technology perception. These findings provide the basis for the next stage of analysis, where the component scores are linked to standardized sales indicators through PCR. Table 8 summarizes these interpretations across both products.

Table 8. Selected principal component interpretations

Product

Principal Component

Dominant Loading Variables

Interpretation Label

iPhone 8

PC1

iOS, USB, Charger, Samsung, Display, Form

Ecosystem compatibility and cross-brand evaluation

iPhone 8

PC11

USB (+), iOS (+), Charger (−), Camera (−), Speed (−)

Technical utility and platform compatibility

Samsung Galaxy Note 7

PC7

Color (+), Android (+), iPhone (−), Camera (+), Wireless (−)

Aesthetic expression and platform identity

Samsung Galaxy Note 7

PC9

Technology (−), Case (−), Battery (+), Display (+)

Risk-sensitive technology and product reliability perception

PC = Principal Component.
4.6 Principal Component Regression Results

After the principal components were obtained and interpreted, the PCR was used to examine the relationship between each component and the standardized success indicator. The component scores were treated as independent variables, while the standardized success indicator was used as the dependent variable. This stage was conducted to identify which latent consumer discussion dimensions were most strongly associated with market success.

4.6.1 Principal Component Regression results for iPhone 8

The PCR results for the iPhone 8 are presented in Table 9. The table shows the correlation coefficient, coefficient of determination, and regression equation for each principal component.

Table 9. PCR summary for iPhone 8

PC

Correlation

R²

Model

1

0.978

0.956

y = 0.4609x + 1.0675

2

0.465

0.216

y = 0.5740x + 0.2693

3

0.309

0.096

y = 0.4483x + 0.3040

4

-0.254

0.065

y = -0.4588x + 0.6481

5

0.210

0.044

y = 0.4451x + 0.3495

6

0.366

0.134

y = 2.4509x + 0.9969

7

0.312

0.097

y = 4.0549x - 0.8270

8

-0.470

0.221

y = -72.598x - 34.251

9

0.192

0.037

y = 25.259x + 14.741

10

0.230

0.053

y = 103.58x - 43.35

11

0.622

0.387

y = 355.36x + 99.366

PCR = Principal Component Regression; PC = Principal Component.

For the iPhone 8, PC1 has the strongest positive association with the standardized success indicator. The correlation value is 0.978, and the R² value is 0.956. This result indicates that PC1 explains 95.6% of the variation in the standardized success indicator. Since PC1 represents a discourse dimension related to ecosystem compatibility, connectivity, display, charging, and cross-brand comparison, the result suggests that these attributes were closely linked to the market performance of the iPhone 8.

This finding is important because it shows that consumer discussions about the iPhone 8 were not only concentrated on individual features. Instead, the strongest market-related dimension combined several connected attributes. Consumers appeared to evaluate the product through the integration of iOS, connectivity, charging support, visual experience, and comparison with Samsung or Android-based products. This pattern is consistent with the positioning of Apple products, where ecosystem integration and user experience often shape consumer perception of product value.

PC11 also shows a meaningful positive relationship with the standardized success indicator. The correlation value is 0.622, and the R² value is 0.387. This means that PC11 explains 38.7% of the variation in the standardized success indicator. Based on the loading structure, PC11 reflects a technical utility dimension that contrasts USB and iOS compatibility with supporting feature performance, such as charger, camera, speed, and Samsung-related comparison. Although its explanatory power is lower than PC1, PC11 still indicates that technical compatibility and functional performance contributed to consumer evaluation of the iPhone 8.

Other components show weaker individual relationships with the standardized success indicator. PC2, PC3, PC5, PC6, PC7, PC9, and PC10 have positive correlations, but their R² values are relatively low. PC4 and PC8 show negative correlations, but their explanatory power is also limited compared with PC1. These results suggest that most individual components outside PC1 and PC11 had weaker roles in explaining iPhone 8 success. Therefore, the market-related discourse for the iPhone 8 was mainly concentrated in one dominant component, supported by one additional technical utility component.

4.6.2 Principal Component Regression results for Samsung Galaxy Note 7

The PCR results for the Samsung Galaxy Note 7 are presented in Table 10. Compared with the iPhone 8, the Note 7 shows a more varied pattern of positive and negative associations across several principal components.

For the Samsung Galaxy Note 7 (Table 10), PC7 has the strongest positive association with the standardized success indicator. The correlation value is 0.943, and the R² value is 0.889. This result indicates that PC7 explains 88.9% of the variation in the standardized success indicator. Based on the loading interpretation, PC7 is associated with color, Android, camera, iPhone, wireless, and case. This component reflects aesthetic expression, Android platform preference, camera-related value, and cross-brand comparison.

The strong positive relationship between PC7 and the success indicator suggests that design-related appeal, platform identity, and camera-related discussion were important positive signals in the Note 7 discourse. Consumers appeared to connect the product with visual differentiation, Android ecosystem preference, and competition with the iPhone. These elements may have supported positive product attention during the early market response.

PC1 also shows a strong positive association with the standardized success indicator, with a correlation value of 0.930 and an R² value of 0.858. This means that PC1 explains 85.8% of the variation in the success indicator. The strength of PC1 indicates that the dominant discourse dimension of the Note 7 was also closely related to market response. This result is consistent with the PCA finding that PC1 captured the largest proportion of variance in Note 7 consumer discussions.

Table 10. PCR summary for Samsung Galaxy Note 7

PC

Correlation

R²

Model

1

0.930

0.858

y = 0.8908x - 0.1744

2

-0.643

0.413

y = -0.3793x + 0.6768

3

-0.304

0.092

y = -0.3540x + 0.2113

4

0.766

0.587

y = 0.7443x + 0.5423

5

-0.138

0.019

y = -0.1773x + 0.2683

6

-0.225

0.051

y = -3.6933x + 2.8782

7

0.943

0.889

y = 93.933x - 11.54

8

0.126

0.016

y = 16.848x - 7.4842

9

-0.915

0.836

y = -146.07x - 63.493

10

-0.089

0.008

y = -12.188x - 2.1697

11

-0.442

0.195

y = -75.467x - 2.6074

PC = Principal Component; PCR = Principal Component Regression.

PC4 shows a moderate to strong positive association, with a correlation value of 0.766 and an R² value of 0.587. Although its explanatory power is lower than PC7 and PC1, PC4 still explains more than half of the variation in the standardized success indicator. This indicates that the Note 7 market response was not explained by only one component. Several consumer discourse dimensions contributed to the observed success pattern.

The most important negative result appears in PC9. This component has a correlation value of -0.915 and an R² value of 0.836. The high R² value shows that PC9 has strong explanatory power, but the negative sign indicates an inverse relationship with the standardized success indicator. This means that when the discourse represented by PC9 increased, the success indicator tended to decrease.

This negative direction is meaningful. PC9 includes technology, case, battery, iPhone, wireless, and display-related signals. Among these attributes, battery is especially important because the Samsung Galaxy Note 7 was historically associated with battery safety issues. Therefore, PC9 can be interpreted as a risk-sensitive discourse dimension. It may represent consumer attention to technology-related concerns, product reliability, battery performance, and comparison with competing products. Unlike PC7, which reflects product appeal, PC9 appears to capture product risk or negative market signals.

Other components show weaker relationships. PC2 has a negative correlation of -0.643 and an R² value of 0.413, suggesting a moderate inverse association. PC3, PC5, PC6, PC8, PC10, and PC11 have relatively lower R² values. These results indicate that their individual contribution to explaining the standardized success indicator was limited compared with PC7, PC1, PC4, and PC9.

4.6.3 Comparative interpretation

The PCR results show different patterns between the two products. For the iPhone 8, market success is mainly explained by PC1, which represents ecosystem compatibility, connectivity, display, and cross-brand comparison. This indicates a more concentrated relationship between consumer discourse and product success.

The iPhone 8 case suggests that a single dominant latent dimension was sufficient to explain most of the variation in the standardized success indicator.

For the Samsung Galaxy Note 7, the relationship is more complex. Several components show strong associations with the standardized success indicator. PC7, PC1, and PC4 show positive relationships, while PC9 shows a strong negative relationship. This pattern suggests that consumer discourse around the Note 7 contained both product appeal and product risk signals. Positive discourse was linked to design, platform identity, camera, and brand comparison. Negative discourse was linked to technology and battery-related concerns.

Overall, the PCR results confirm that social media-derived principal components can be used to examine the relationship between consumer attribute discussions and product success. The findings also show that high consumer attention does not always imply positive market value. Some discussion dimensions may signal consumer interest, while others may indicate risk, dissatisfaction, or potential product failure. This distinction is important for product development because firms need to identify not only which attributes are frequently discussed, but also whether those discussions are positively or negatively associated with market outcomes.

These results should be interpreted as associative rather than causal. The regression analysis shows the strength and direction of the relationship between social media-derived components and the standardized success indicator. It does not prove that consumer discussions directly caused product success or failure. However, the results provide useful empirical evidence that online consumer discourse contains measurable signals related to market response.

4.7 Product Development Insights

The findings provide several implications for product development. First, social media discussions can help identify product attributes that receive strong consumer attention during and after product launch. Across the two smartphone cases, consumers repeatedly discussed display quality, camera capability, battery performance, charging system, operating system, wireless features, design, and brand comparison. These attributes represent visible areas of consumer evaluation. They should be monitored closely because they can shape early market response and influence how consumers judge product value.

Second, the PCA results show that consumers do not evaluate product attributes as separate features. Instead, attributes tend to form broader perception dimensions. In the iPhone 8 case, the most influential dimension was related to ecosystem compatibility, connectivity, visual experience, charging support, and competitive comparison. This finding suggests that the market value of the iPhone 8 was strongly associated with the integration of hardware, software, accessories, and the Apple ecosystem. For product managers, this means that product success may depend not only on improving individual features, but also on ensuring that these features work together as a coherent user experience.

Third, the Samsung Galaxy Note 7 case shows that social media discourse can capture both product appeal and product risk. Positive components were associated with design, color, Android platform identity, camera, and comparison with the iPhone. These dimensions reflect the product’s appeal as a premium Android smartphone. However, the negative component was linked to technology and battery-related signals. This pattern is important because the Galaxy Note 7 was historically associated with battery safety issues. The result suggests that social media analytics can help firms detect early warning signals when consumer discussions begin to concentrate around risk-sensitive attributes.

Fourth, the PCR results show that the proposed framework extends beyond descriptive text mining. It does not only identify which attributes consumers discuss most often. It also examines which latent discussion dimensions are associated with market success. This distinction is important for product development. A frequently discussed attribute may reflect interest, dissatisfaction, comparison, or concern. By linking social media-derived components to standardized success indicators, firms can better distinguish between attributes that support market performance and attributes that indicate potential product risk.

In practice, firms can apply the proposed framework at three distinct stages of the product decision process. During product planning, the text mining phase can be used to audit the competitive attribute landscape before a product launch. By identifying the dominant attributes in consumer discourse for competing products, the PCA can inform feature investment priorities and help predict consumer expectations. During launch evaluation, the PCA results can be used to monitor whether consumer discussion is coalescing around positive value dimensions or risk-sensitive attributes. A shift in dominant component scores towards risk-related dimensions can trigger the need for early corrective communication or product intervention. During attribute prioritization, the PCR results can be used to rank which consumer discussion dimensions are most strongly associated with market performance, so that product teams can allocate resources to the attributes and feature clusters that have the highest market relevance. These three application stages make the framework operational beyond academic analysis, and suitable for integration into regular product intelligence workflows.

4.8 Discussion of the Proposed Framework

The findings support the relevance of a data-driven social media framework for product development. The proposed framework shows how unstructured consumer discourse can be transformed into measurable product development insights. Text mining first identified the dominant product attributes discussed by consumers. PCA then reduced overlapping and correlated attributes into interpretable latent components. PCR finally examined the relationship between these components and standardized sales indicators. This sequence provides a systematic analytical path from online consumer discussion to market success analysis.

The results extend prior studies that used social media analytics and online consumer data for product development. Previous research has shown that text mining can identify product attributes, customer opinions, and improvement opportunities from online reviews or social media data. For example, Zhang et al. (2018) used online review mining to support product innovation, while Hong & Wang (2021) applied topic mining and deep learning to summarize customer opinions. Rathore & Ilavarasan (2020) also showed that Twitter/X analytics can capture consumer emotions before and after product launch. These studies confirm that online consumer data is useful for understanding customer response. However, most of them focus on extracting topics, opinions, or sentiments. They do not fully examine how these attributes form latent structures or how those structures relate to market success.

This study addresses that limitation by adding PCA and PCR after the text mining stage. PCA is important because consumer-discussed attributes are often interrelated. For example, smartphone users do not discuss display, camera, battery, charger, operating system, and brand comparison as isolated features. These attributes form broader perception dimensions. In the iPhone 8 case, the dominant latent dimension was related to ecosystem compatibility, connectivity, display, charging support, and cross-brand comparison. This result suggests that the market value of the iPhone 8 was strongly connected to Apple’s integrated ecosystem and user experience. The finding is consistent with the view that product value in technology markets depends not only on individual features, but also on how features work together as a coherent product experience.

The results also contribute to studies on product opportunity mining and complaint-based product improvement. Wang et al. (2023) used social media analytics to mine customer complaints and identify product opportunities. Zhang & Song (2024) combined text mining and decision-making methods to support product improvement in a big data environment. These studies show that consumer-generated text can help firms identify improvement priorities. However, product opportunities and complaints do not always indicate market success. A frequently discussed attribute may reflect satisfaction, dissatisfaction, curiosity, or risk. The PCR results in this study help distinguish these meanings by linking latent components to standardized sales indicators.

The comparison between iPhone 8 and Samsung Galaxy Note 7 shows that product context matters. The iPhone 8 case shows a more positive and concentrated market-related discourse. PC1 produced the strongest relationship with the success indicator, with an R2 value of 0.956. This indicates that the main discourse dimension captured most of the market-related variation. In contrast, the Galaxy Note 7 case showed both positive and negative market-related signals. Some components were positively associated with the success indicator, such as components related to aesthetic appeal, Android identity, camera, and competitive comparison. However, other components showed negative associations, especially those linked to technology and battery-related concerns. This pattern reflects the historical nature of the Galaxy Note 7 case, where strong product appeal existed together with serious product risk.

This finding is important for social media-based product development. High consumer attention does not always represent positive product value. In some cases, high discussion intensity may signal dissatisfaction, safety concerns, or product failure. This is especially relevant for the Galaxy Note 7, where battery-related discourse became a risk-sensitive signal. Therefore, firms should not only monitor the volume of consumer discussions. They also need to examine the structure and direction of those discussions in relation to market outcomes.

The study also extends research on data-driven product planning. Panzner et al. (2024) emphasized the need for structured data analytics pipelines that connect data sources, analytical methods, and product planning decisions. The present framework responds to that need by combining text mining, PCA, and PCR in one analytical process. It does not stop at identifying what consumers discuss. It also identifies which latent consumer perception dimensions have stronger market relevance. This makes the framework more useful for product managers who need to prioritize attributes, evaluate launch response, and detect early product risks.

Overall, the proposed framework can support product development in four ways. First, it helps firms identify product attributes that receive strong consumer attention from social media data. Second, it reduces overlapping attributes into latent consumer perception dimensions, which improves interpretability and reduces redundancy. Third, it links these dimensions to market success indicators, allowing firms to identify which consumer discussions matter most for product performance. Fourth, it supports product risk monitoring by distinguishing positive market signals from negative or risk-sensitive discourse.

These findings show that social media analytics can move beyond descriptive insight extraction. When combined with PCA and PCR, social media data can provide empirical signals for product development and market success analysis. The framework can support product planning, launch monitoring, competitive benchmarking, and early corrective action when consumer discourse begins to concentrate around critical product attributes.

5. Conclusion

This study proposed a data-driven social media framework for product innovation by integrating text mining, PCA, and PCR. The framework was designed to transform unstructured consumer discussions from social media into measurable product innovation insights and to examine their relationship with market success. Using Apple iPhone 8 and Samsung Galaxy Note 7 as empirical cases, this study shows that social media discourse can capture consumer attention toward product attributes, platform comparison, and potential product risks.

The text mining results indicate that consumers discussed several key smartphone attributes across both products, including visual features, operating system, connectivity, battery, charging, camera, platform identity, and cross-brand comparison. These findings confirm that consumer evaluation of smartphones is multidimensional. Consumers do not only assess technical specifications, but also consider ecosystem compatibility, brand identity, daily usability, and product comparison with competing alternatives.

The PCR results further show that selected principal components are strongly associated with standardized sales indicators. For the iPhone 8, PC1 has the strongest relationship with standardized sales, with an R2 value of 0.956. This suggests that the iPhone 8 market response was strongly linked to consumer discussions related to ecosystem compatibility, connectivity, visual experience, charging support, and cross-brand comparison. For the Samsung Galaxy Note 7, PC7 shows the strongest relationship with standardized sales, with an R2value of 0.889. Other components show both positive and negative relationships, indicating that social media discourse can reflect not only product appeal, but also product risk.

The main contribution of this study lies in its integrated analytical framework. Unlike studies that focus only on topic extraction, sentiment analysis, or complaint mining, this study links consumer attribute discussions to market success indicators. This approach allows firms to identify which latent dimensions of consumer discourse have stronger market relevance. It also helps distinguish between attributes that support product value and attributes that may signal consumer concern or product risk.

From a managerial perspective, the proposed framework can support product innovation in several ways. It can help firms identify consumer-prioritized attributes from social media data, reduce overlapping product attributes into interpretable components, monitor product launch responses, compare competing products, and detect early warning signals from consumer discourse. These insights can support evidence-based product development and help firms prioritize innovation efforts based on market-relevant consumer signals.

Despite these contributions, several limitations should be acknowledged. The framework was only tested on two smartphone products. Therefore, its generalizability to other product categories (e.g., consumer electronics, fast-moving consumer goods, and industrial products) needs to be tested, as different product types may have different discourse patterns, which may affect the structure of the components. The study also relied on data from a single collection period and a single social media platform because consumer discourse evolves across the product lifecycle and varies by platform, future research could apply the framework to time-series data and compare results across platforms such as Twitter/X, Reddit, and product review sites. Finally, future validation across non-English language markets and emerging product categories would further strengthen the framework's external validity.

Author Contributions

Conceptualization, D.A.P. and R.D.K.; methodology, D.A.P.; software, R.D.K.; validation, D.A.P. and R.D.K.; formal analysis, D.A.P.; investigation, R.D.K.; resources, D.A.P.; data curation, D.A.P.; writing—original draft preparation, D.A.P. and R.D.K.; writing—review and editing, D.A.P. and R.D.K.; visualization, D.A.P.; supervision, R.D.K.; project administration, D.A.P.; funding acquisition, R.D.K. All authors have read and agreed to the published version of the manuscript.

Data Availability

All relevant data are included within the article, and additional data can be provided by the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References
Cai, M. & Zhang, X. (2025). Consumer preference analysis integrating online reviews: A multiple criteria group approach considering individual stochastic behavior. 4OR, 23(1), 97–122. [Crossref]
Cooper, R. G. (2019). The drivers of success in new-product development. Ind. Mark. Manag., 76, 36–47. [Crossref]
Cui, A. S. & Wu, F. (2016). Utilizing customer knowledge in innovation: Antecedents and impact of customer involvement on new product performance. J. Acad. Mark. Sci., 44(4), 516–538. [Crossref]
Ding, Y., Wu, P., Zhao, J., & Zhou, L. (2025). Forecasting product sales using text mining: A case study in new energy vehicle. Electron. Commer. Res., 25(1), 495–527. [Crossref]
Dissanayake, R. & Amarasuriya, T. (2015). Role of brand identity in developing global brands: A literature-based review on case comparison between apple iPhone vs Samsung smartphone brands. Res. J. Bus. Manag., 2(3), 430–440.
Geissinger, A., Laurell, C., Öberg, C., & Sandström, C. (2023). Social media analytics for innovation management research: A systematic literature review and future research agenda. Technovation, 123, 102712. [Crossref]
Greenacre, M., Groenen, P. J., Hastie, T., d’Enza, A. I., Markos, A., & Tuzhilina, E. (2022). Principal component analysis. Nat. Rev. Methods Primers, 2(1), 100. [Crossref]
Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2019). Multivariate Data Analysis. Cengage Learning.
Han, R., Brennecke, J., Borah, D., & Lam, H. K. S. (2025). The use of social media in different phases of the new product development process: A systematic literature review. R D Manag., 55(1), 108–126. [Crossref]
Harvey, S. & Kou, C.Y. (2013). Collective engagement in creative tasks: The role of evaluation in the creative process in groups. Adm. Sci. Q., 58(3), 346–386. [Crossref]
Hong, M. & Wang, H. (2021). Research on customer opinion summarization using topic mining and deep neural network. Math. Comput. Simul., 185, 88–114. [Crossref]
Klein, M. A. & Şener, F. (2023). Product innovation, diffusion and endogenous growth. Rev. Econ. Dyn., 48, 178–201. [Crossref]
Li, S., Liu, F., Zhang, Y., Zhu, B., Zhu, H., & Yu, Z. (2022). Text mining of user-generated content (UGC) for business applications in e-commerce: A systematic review. Mathematics, 10(19), 3554. [Crossref]
Lin, K. Y. (2018). User experience-based product design for smart production to empower industry 4.0 in the glass recycling circular economy. Comput. Ind. Eng., 125, 729–738. [Crossref]
Morgan, T., Obal, M., & Anokhin, S. (2018). Customer participation and new product performance: Towards the understanding of the mechanisms and key contingencies. Res. Policy, 47(2), 498–510. [Crossref]
Najafi-Tavani, Z., Mousavi, S., Zaefarian, G., & Naudé, P. (2020). Relationship learning and international customer involvement in new product design: The moderating roles of customer dependence and cultural distance. J. Bus. Res., 120, 42–58. [Crossref]
Nasrabadi, M. A., Beauregard, Y., & Ekhlassi, A. (2024). The implication of user-generated content in new product development process: A systematic literature review and future research agenda. Technol. Forecast. Soc. Change, 206, 123551. [Crossref]
Panzner, M., von Enzberg, S., & Dumitrescu, R. (2024). Developing a data analytics toolbox for data-driven product planning: A review and survey methodology. AI EDAM, 38, e18. [Crossref]
Purnama, D. A., Subagyo, & Masruroh, N. A. (2023). Online data-driven concurrent product-process-supply chain design in the early stage of new product development. J. Open Innov. Technol. Mark. Complex., 9(3), 100093. [Crossref]
Rathore, A. K. & Ilavarasan, P. V. (2020). Pre-and post-launch emotions in new product development: Insights from twitter analytics of three products. Int. J. Inf. Manage., 50, 111–127. [Crossref]
Sońta-Drączkowska, E., Cichosz, M., Klimas, P., & Pilewicz, T. (2025). Co-creating innovations with users: A systematic literature review and future research agenda for project management. Eur. Manag. J., 43(2), 321–339. [Crossref]
Tian, Q., Cao, G., & Weerawardena, J. (2024). Strategic use of social media in new product development in B2B firms: The role of absorptive capacity. Ind. Mark. Manag., 120, 132–145. [Crossref]
Ulrich, K. T., Eppinger, S. D., & Yang, M. C. (2008). Product Design and Development. McGraw-Hill Education.
Wang, J., Lai, J. Y., & Lin, Y. H. (2023). Social media analytics for mining customer complaints to explore product opportunities. Comput. Ind. Eng., 178, 109104. [Crossref]
Wu, P., Tang, T., Zhou, L., & Martínez, L. (2024). A decision-support model through online reviews: Consumer preference analysis and product ranking. Inf. Process. Manag., 61(4), 103728. [Crossref]
Xie, X. & Jia, Y. (2016). Consumer involvement in new product development: A case study from the online virtual community. Psychol. Mark., 33(12), 1187–1194. [Crossref]
Yakubu, H. & Kwong, C. K. (2021). Forecasting the importance of product attributes using online customer reviews and Google Trends. Technol. Forecast. Soc. Change, 171, 120983. [Crossref]
Zhang, F. & Song, W. (2024). Product improvement in a big data environment: A novel method based on text mining and large group decision making. Expert Syst. Appl., 245, 123015. [Crossref]
Zhang, H., Rao, H., & Feng, J. (2018). Product innovation based on online review data mining: A case study of Huawei phones. Electron. Commer. Res., 18(1), 3–22. [Crossref]

Cite this:
APA Style
IEEE Style
BibTex Style
MLA Style
Chicago Style
GB-T-7714-2015
Purnama, D. A. & Kurnia, R. D. (2026). A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success. J. Res. Innov. Technol., 5(2), 223-241. https://doi.org/10.56578/jorit050206
D. A. Purnama and R. D. Kurnia, "A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success," J. Res. Innov. Technol., vol. 5, no. 2, pp. 223-241, 2026. https://doi.org/10.56578/jorit050206
@research-article{Purnama2026ADS,
title={A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success},
author={Dwi Adi Purnama and Ratih Dianingtyas Kurnia},
journal={Journal of Research, Innovation and Technologies},
year={2026},
page={223-241},
doi={https://doi.org/10.56578/jorit050206}
}
Dwi Adi Purnama, et al. "A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success." Journal of Research, Innovation and Technologies, v 5, pp 223-241. doi: https://doi.org/10.56578/jorit050206
Dwi Adi Purnama and Ratih Dianingtyas Kurnia. "A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success." Journal of Research, Innovation and Technologies, 5, (2026): 223-241. doi: https://doi.org/10.56578/jorit050206
PURNAMA D A, KURNIA R D. A Data-Driven Social Media Framework for Product Innovation: Linking Consumer Attribute Discussions to Market Success[J]. Journal of Research, Innovation and Technologies, 2026, 5(2): 223-241. https://doi.org/10.56578/jorit050206
cc
©2026 by the author(s). Published by Acadlore Publishing Services Limited, Hong Kong. This article is available for free download and can be reused and cited, provided that the original published version is credited, under the CC BY 4.0 license.