<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>tfidf Archives - Turbolab Technologies</title>
	<atom:link href="https://turbolab.in/tag/tfidf/feed/" rel="self" type="application/rss+xml" />
	<link>https://turbolab.in/tag/tfidf/</link>
	<description>Big Data and News Analysis Startup in Kochi</description>
	<lastBuildDate>Fri, 03 Dec 2021 05:38:55 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.2</generator>

<image>
	<url>https://i0.wp.com/turbolab.in/wp-content/uploads/2018/03/turbo_black_trans-space.png?fit=32%2C32&#038;ssl=1</url>
	<title>tfidf Archives - Turbolab Technologies</title>
	<link>https://turbolab.in/tag/tfidf/</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">98237731</site>	<item>
		<title>Feature Extraction in Natural Language Processing</title>
		<link>https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/</link>
					<comments>https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/#respond</comments>
		
		<dc:creator><![CDATA[Anthony]]></dc:creator>
		<pubDate>Fri, 08 Oct 2021 10:49:26 +0000</pubDate>
				<category><![CDATA[Technology]]></category>
		<category><![CDATA[bag of words]]></category>
		<category><![CDATA[Data Science]]></category>
		<category><![CDATA[feature extraction]]></category>
		<category><![CDATA[machine learning]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[nlp]]></category>
		<category><![CDATA[tfidf]]></category>
		<category><![CDATA[word embeddings]]></category>
		<category><![CDATA[word2vec]]></category>
		<guid isPermaLink="false">https://turbolab.in/?p=629</guid>

					<description><![CDATA[<p>In simple terms, Feature Extraction is transforming textual data into numerical data. In Natural Language Processing, Feature Extraction is a very trivial method to be followed to better understand the context. After cleaning and normalizing textual data, we need to transform it into their features for modeling, as the machine does not compute textual data. [&#8230;]</p>
<p>The post <a href="https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/">Feature Extraction in Natural Language Processing</a> appeared first on <a href="https://turbolab.in">Turbolab Technologies</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><span style="font-weight: 400;">In simple terms, Feature Extraction is transforming textual data into numerical data. In Natural Language Processing, Feature Extraction is a very trivial method to be followed to better understand the context. After cleaning and normalizing textual data, we need to transform it into their features for modeling, as the machine does not compute textual data. So we go for numerical representation for individual words as it’s easy for the computer to process numbers.</span></p>
<p><span style="font-weight: 400;">In this blog, we will discuss various feature extraction methods with examples using sklearn and gensim.</span></p>
<p>&nbsp;</p>
<ul>
<li><b>Countvectorizer</b></li>
</ul>
<ul>
<li><strong>TF-IDF Vectorizer</strong></li>
</ul>
<ul>
<li><strong>Word Embeddings</strong></li>
</ul>
<p>&nbsp;</p>
<h3></h3>
<h2><strong>Countvectorizer</strong></h2>
<p>&nbsp;</p>
<p><img data-recalc-dims="1" fetchpriority="high" decoding="async" data-attachment-id="632" data-permalink="https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/countvectorizer-2/" data-orig-file="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?fit=1726%2C800&amp;ssl=1" data-orig-size="1726,800" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;}" data-image-title="countvectorizer" data-image-description="" data-image-caption="" data-large-file="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?fit=800%2C371&amp;ssl=1" class="alignnone wp-image-632" src="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=642%2C297&#038;ssl=1" alt="" width="642" height="297" srcset="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=300%2C139&amp;ssl=1 300w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=768%2C356&amp;ssl=1 768w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=1024%2C475&amp;ssl=1 1024w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=1080%2C501&amp;ssl=1 1080w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=1280%2C593&amp;ssl=1 1280w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=980%2C454&amp;ssl=1 980w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?resize=480%2C222&amp;ssl=1 480w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?w=1726&amp;ssl=1 1726w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/countvectorizer-1.png?w=1600&amp;ssl=1 1600w" sizes="(max-width: 642px) 100vw, 642px" /></p>
<p>&nbsp;</p>
<p><span style="font-weight: 400;">It is a simple and flexible way of extracting features from documents. A Countvectorizer model is a representation of text that describes the occurrence of words within a document. We just keep track of word counts and disregard the grammatical details and the word order. It is called a “bag of words” because any information about the order or structure of words in the document is discarded. The model is only concerned with whether known words occur in the document, not wherein the document.</span></p>
<p><span style="font-weight: 400;">Here is a basic snippet of using count vectorization to get vectors</span></p>
<p>&nbsp;</p>
<blockquote><p><strong><i>from sklearn.feature_extraction.text import CountVectorizer</i></strong></p>
<p>&nbsp;</p>
<p><strong><i>corpus = [&#8220;We become what we think about&#8221;, &#8220;Happiness is not something readymade. It comes from your own actions&#8221;]</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># initialize count vectorizer object</i></strong></p>
<p><strong><i>vect = CountVectorizer()</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># get counts of each token (word) in text data</i></strong></p>
<p><strong><i>X = vect.fit_transform(corpus)</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># convert sparse matrix to numpy array to view</i></strong></p>
<p><strong><i>X.toarray()</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># view token vocabulary and counts</i></strong></p>
<p><strong><i>print(&#8220;vocabulary&#8221;, vect.vocabulary_)</i></strong></p>
<p><strong><i>print(&#8220;shape&#8221;, X.shape)</i></strong></p>
<p><strong><i>print(&#8216;vectors: &#8216;, X.toarray())</i></strong></p></blockquote>
<p>&nbsp;</p>
<h4><strong>Output</strong></h4>
<p>&nbsp;</p>
<blockquote><p><em><span style="font-weight: 400;"><strong>Vocabulary</strong> :  {&#8216;we&#8217;: 8, &#8216;become&#8217;: 1, &#8216;what&#8217;: 9, &#8216;think&#8217;: 7, &#8216;about&#8217;: 0, &#8216;happiness&#8217;: 2, &#8216;is&#8217;: 3, &#8216;not&#8217;: 4, &#8216;something&#8217;: 6, &#8216;readymade&#8217;: 5}</span></em></p>
<p>&nbsp;</p>
<p><em><span style="font-weight: 400;"><strong>Shape</strong> :  (2, 10)</span></em></p>
<p>&nbsp;</p>
<p><em><span style="font-weight: 400;"><strong>Vectors</strong> :  [[1 1 0 0 0 0 0 1 2 1]</span></em></p>
<p><em><span style="font-weight: 400;">                 [0 0 1 1 1 1 1 0 0 0]]</span></em></p></blockquote>
<p>&nbsp;</p>
<p>&nbsp;</p>
<p>&nbsp;</p>
<h3></h3>
<h2><b>TF &#8211; IDF Vectorizer (Term Frequency &#8211; Inverse Document Frequency)</b></h2>
<p>&nbsp;</p>
<p><img data-recalc-dims="1" decoding="async" data-attachment-id="642" data-permalink="https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/tf-idf/" data-orig-file="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?fit=1200%2C399&amp;ssl=1" data-orig-size="1200,399" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;}" data-image-title="tf-idf" data-image-description="" data-image-caption="" data-large-file="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?fit=800%2C266&amp;ssl=1" class="alignnone wp-image-642" src="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?resize=705%2C235&#038;ssl=1" alt="" width="705" height="235" srcset="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?resize=300%2C100&amp;ssl=1 300w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?resize=768%2C255&amp;ssl=1 768w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?resize=1024%2C340&amp;ssl=1 1024w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?resize=1080%2C359&amp;ssl=1 1080w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?resize=980%2C326&amp;ssl=1 980w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?resize=480%2C160&amp;ssl=1 480w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/tf-idf.png?w=1200&amp;ssl=1 1200w" sizes="(max-width: 705px) 100vw, 705px" /></p>
<p>&nbsp;</p>
<p><span style="font-weight: 400;">TF-IDF is short for term frequency-inverse document frequency. It’s designed to reflect how important a word is to a document in a collection or corpus.</span></p>
<p><span style="font-weight: 400;">The TF-IDF value increases proportionally to the number of times a word appears in the document and is offset by the number of documents in the corpus that contain the word, which helps to adjust for the fact that some words appear more frequently in general.</span></p>
<p><span style="font-weight: 400;">And similar to the Countvectorizer, </span><i><span style="font-weight: 400;">sklearn.feature_extraction.text</span></i><span style="font-weight: 400;"> provides a method.</span></p>
<p>&nbsp;</p>
<blockquote><p><strong><i>from sklearn.feature_extraction.text import TfidfVectorizer</i></strong></p>
<p>&nbsp;</p>
<p><strong><i>corpus = [&#8220;We become what we think about&#8221;, &#8220;Happiness is not something readymade.&#8221;]</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># initialize tf-idf vectorizer object</i></strong></p>
<p><strong><i>vectorizer = TfidfVectorizer()</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># compute bag of word counts and tf-idf values</i></strong></p>
<p><strong><i>tf = vectorizer.fit_transform(corpus)</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># convert sparse matrix to numpy array to view</i></strong></p>
<p><strong><i>print(&#8220;Vocabulary&#8221;, vectorizer.vocabulary_)</i></strong></p>
<p><strong><i>print(&#8220;idf&#8221;, vectorizer.idf_)</i></strong></p>
<p><strong><i>print(&#8220;Vectors&#8221;, tf.toarray())</i></strong></p></blockquote>
<p>&nbsp;</p>
<h4><b>Output:</b></h4>
<p>&nbsp;</p>
<blockquote><p><em><span style="font-weight: 400;"><strong>Vocabulary</strong> : {&#8216;we&#8217;: 8, &#8216;become&#8217;: 1, &#8216;what&#8217;: 9, &#8216;think&#8217;: 7, &#8216;about&#8217;: 0, &#8216;happiness&#8217;: 2, &#8216;is&#8217;: 3, &#8216;not&#8217;: 4, &#8216;something&#8217;: 6, &#8216;readymade&#8217;: 5}</span></em></p>
<p>&nbsp;</p>
<p><em><span style="font-weight: 400;"><strong>idf</strong> : [1.40546511 1.40546511 1.40546511 1.40546511 1.40546511 1.40546511</span></em></p>
<p><em><span style="font-weight: 400;"> 1.40546511 1.40546511 1.40546511 1.40546511]</span></em></p>
<p>&nbsp;</p>
<p><em><span style="font-weight: 400;"><strong>Vectors</strong> : [[0.35355339 0.35355339 0.         0.         0.         0.</span></em></p>
<ol>
<li><em><span style="font-weight: 400;">         0.35355339 0.70710678 0.35355339]</span></em></li>
</ol>
<p><em><span style="font-weight: 400;">      [0.         0.         0.4472136  0.4472136  0.4472136  0.4472136</span></em></p>
<p><em><span style="font-weight: 400;">       0.4472136  0.         0.         0.        ]]</span></em></p></blockquote>
<p>&nbsp;</p>
<p>&nbsp;</p>
<h2><b>Word Embeddings</b></h2>
<p>&nbsp;</p>
<p><span style="font-weight: 400;">Word embedding is a learned representation of text, where each word is represented as a real-valued vector in a lower-dimensional space.</span></p>
<p><span style="font-weight: 400;">In simple terms, word embeddings are the texts converted into numbers and there may be different numerical representations of the same text, but texts with similar context have similar representations.</span></p>
<p><span style="font-weight: 400;">Word embedding preserves contexts and relationships of words so that it detects similar words more accurately.</span></p>
<p><span style="font-weight: 400;">Word embedding has several different implementations such as word2vec, GloVe, FastText etc.</span></p>
<p><span style="font-weight: 400;">Here we will explain word2vec, as it is the most popular implementation.</span></p>
<p>&nbsp;</p>
<h3><b>Word2vec</b></h3>
<p>&nbsp;</p>
<p><span style="font-weight: 400;">Word2vec is widely used in most of the NLP models. It transforms every word into vectors. Word2vec can make the most accurate predictions about the meaning of words. It can capture the contextual meaning of words very well. Word vectors are positioned in the vector space such that words that share common contexts in the corpus are located close to one another in space.</span></p>
<p><span style="font-weight: 400;">There are two neural embedding algorithms:</span></p>
<p>&nbsp;</p>
<ul>
<li><b>Continuous Bag-of-Words (CBOW) &#8211; <span style="font-weight: 400;">predicts target word from context</span></b></li>
<li><b>Skip-gram &#8211; </b>predicts context from the target word</li>
</ul>
<p>&nbsp;</p>
<p><img data-recalc-dims="1" decoding="async" data-attachment-id="638" data-permalink="https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/word2vec_auto_x2/" data-orig-file="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?fit=3248%2C1228&amp;ssl=1" data-orig-size="3248,1228" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;1&quot;}" data-image-title="word2vec_auto_x2" data-image-description="" data-image-caption="" data-large-file="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?fit=800%2C302&amp;ssl=1" class="alignnone wp-image-638" src="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=669%2C252&#038;ssl=1" alt="" width="669" height="252" srcset="https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=300%2C113&amp;ssl=1 300w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=768%2C290&amp;ssl=1 768w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=1024%2C387&amp;ssl=1 1024w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=1080%2C408&amp;ssl=1 1080w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=1280%2C484&amp;ssl=1 1280w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=980%2C371&amp;ssl=1 980w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?resize=480%2C181&amp;ssl=1 480w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?w=1600&amp;ssl=1 1600w, https://i0.wp.com/turbolab.in/wp-content/uploads/2021/10/word2vec_auto_x2.jpg?w=2400&amp;ssl=1 2400w" sizes="(max-width: 669px) 100vw, 669px" /></p>
<p>&nbsp;</p>
<p><span style="font-weight: 400;">Here is an example of Word2vec using Gensim. Gensim is a python library for NLP.</span></p>
<p>&nbsp;</p>
<blockquote><p><strong><i>from gensim.models import Word2Vec</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># Get document data.</i></strong></p>
<p><strong><i>common_texts = [[&#8216;interface&#8217;, &#8216;computer&#8217;, &#8216;technology&#8217;],</i></strong></p>
<p><strong><i> [&#8216;survey&#8217;, &#8216;computer&#8217;, &#8216;system&#8217;, &#8216;response&#8217;],</i></strong></p>
<p><strong><i> [ &#8216;brother&#8217;, &#8216;boy&#8217;, &#8216;man&#8217;, &#8216;animal&#8217;, &#8216;human&#8217;]]</i></strong></p>
<p>&nbsp;</p>
<p><strong><i># Initializing Model</i></strong></p>
<p><strong><i>model = Word2Vec(common_texts, window=5, min_count=1, workers=4)</i></strong></p></blockquote>
<p>&nbsp;</p>
<h3><strong>Result 1 :</strong></h3>
<p>&nbsp;</p>
<p><b># Get most similar words of &#8220;computer&#8221;</b></p>
<p><b>model.wv.most_similar(&#8220;computer&#8221;)</b></p>
<p>&nbsp;</p>
<h4><b>Output :</b><span style="font-weight: 400;"> </span></h4>
<p>&nbsp;</p>
<blockquote><p><span style="font-weight: 400;">[(&#8216;technology&#8217;, 0.21617145836353302),</span></p>
<p><span style="font-weight: 400;"> (&#8216;system&#8217;, 0.09291724115610123),</span></p>
<p><span style="font-weight: 400;"> (&#8216;interface&#8217;, 0.06285080313682556),</span></p>
<p><span style="font-weight: 400;"> (&#8216;survey&#8217;, 0.027057476341724396),</span></p>
<p><span style="font-weight: 400;"> (&#8216;response&#8217;, 0.016134709119796753),</span></p>
<p><span style="font-weight: 400;"> (&#8216;human&#8217;, -0.010839173570275307),</span></p>
<p><span style="font-weight: 400;"> (&#8216;boy&#8217;, -0.02775038219988346),</span></p>
<p><span style="font-weight: 400;"> (&#8216;animal&#8217;, -0.052346907556056976),</span></p>
<p><span style="font-weight: 400;"> (&#8216;brother&#8217;, -0.05987627059221268),</span></p>
<p><span style="font-weight: 400;"> (&#8216;man&#8217;, -0.111670583486557)]</span></p></blockquote>
<p>&nbsp;</p>
<h3><b>Result 2 :</b></h3>
<p>&nbsp;</p>
<p><b># Get most similar words of &#8220;computer&#8221;</b></p>
<p><b>model.wv.most_similar(&#8220;human&#8221;)</b></p>
<p>&nbsp;</p>
<h4><b>Output :</b></h4>
<p>&nbsp;</p>
<blockquote><p><span style="font-weight: 400;">[(&#8216;man&#8217;, 0.0679759532213211),</span></p>
<p><span style="font-weight: 400;"> (&#8216;survey&#8217;, 0.03364055976271629),</span></p>
<p><span style="font-weight: 400;"> (&#8216;brother&#8217;, 0.00939119141548872),</span></p>
<p><span style="font-weight: 400;"> (&#8216;boy&#8217;, 0.004503018222749233),</span></p>
<p><span style="font-weight: 400;"> (&#8216;computer&#8217;, -0.010839177295565605),</span></p>
<p><span style="font-weight: 400;"> (&#8216;animal&#8217;, -0.02365921437740326),</span></p>
<p><span style="font-weight: 400;"> (&#8216;technology&#8217;, -0.09575347602367401),</span></p>
<p><span style="font-weight: 400;"> (&#8216;response&#8217;, -0.11410721391439438),</span></p>
<p><span style="font-weight: 400;"> (&#8216;system&#8217;, -0.11555543541908264),</span></p>
<p><span style="font-weight: 400;"> (&#8216;interface&#8217;, -0.13429945707321167)]</span></p></blockquote>
<p>&nbsp;</p>
<h2><b>Conclusion</b></h2>
<p>&nbsp;</p>
<p><span style="font-weight: 400;">In this post, we have discovered different types of text Feature Extraction Methods where we moved from non-context vectorization methods (count vectorizer/BOWs) to context preserving methods (TF-IDF/Word Embeddings). We have explored the above methods practically using Scikit-learn (sklearn) and Gensim libraries.</span></p>
<p><span style="font-weight: 400;">There are other advanced techniques for Word Embeddings like Facebook&#8217;s FastText. We will discuss them in our coming blogs.</span></p>
<p><span style="font-weight: 400;">Apart from Word Embeddings, Dimension Reductionality is also a Feature Extraction technique that aims to reduce the number of features in a dataset by creating new features from the existing ones and then discarding the original features.</span></p>
<p><span style="font-weight: 400;">Different techniques that you can explore for dimension reductional are Principal Components Analysis (PCA), Linear Discriminant Analysis (LDA), t-distributed Stochastic Neighbor Embedding (t-SNE), and many more.</span></p>
<p>The post <a href="https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/">Feature Extraction in Natural Language Processing</a> appeared first on <a href="https://turbolab.in">Turbolab Technologies</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://turbolab.in/feature-extraction-in-natural-language-processing-nlp/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">629</post-id>	</item>
	</channel>
</rss>
