PageSourceSearch

https://www.kinetica.com/assets/bring-deep-learning-model-kinetica-f8L8iTtv.js

js kinetica.com collected 2026-10-02 06:35:42 UTC 14,820 bytes, 82 lines download raw bytes

1const e="bring-deep-learning-model-kinetica",t="Bring Your Deep Learning Model to Kinetica",a="How can we avoid the data science black hole of complexity, unpredictability, and disastrous failures and actually make it work for our organizations? According to this recent Eckerson article the key is operationalizing data science: We, as a field, and I mean academics, scientists, product developers, data scientists, consultants … everybody … need to redirect our efforts towards operationalizing data science. We as practitioners can unleash the power of data science only when we make it safe and find a way to fit it into normal business processes. Here at Kinetica we couldn’t agree more! What’s the point of building a brilliant model if you can’t actually get it into production? That’s the problem we’re helping to solve by streamlining this machine learning (ML) operationalization process. Kinetica’s GPU-accelerated engine can easily and quickly run deep learning models in parallel across multiple machines. Our CUDA-optimized user-defined functions API, makes it simple to prepare, train and deploy predictive analytics and ML. In this blog, we’ll go through the process of […]",n="2018-08-31",s="Zhe Wu",o="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im3-1.png",i=["Developer Blog"],r=`<p class="text-gray-600 leading-relaxed mb-4">How can we avoid the data science black hole of complexity, unpredictability, and disastrous failures and actually make it work for our organizations?</p>
2<p class="text-gray-600 leading-relaxed mb-4">According to this <a href="https://www.eckerson.com/articles/operationalizing-data-science-13-challenges" class="text-kinetica-600 hover:underline">recent Eckerson article</a> the key is operationalizing data science:</p>
3<blockquote class="border-l-4 border-kinetica-600 pl-4 italic text-gray-600 mb-4">
4  <p>We, as a field, and I mean academics, scientists, product developers, data scientists, consultants … everybody … <strong>need to redirect our efforts towards operationalizing data science.</strong> We as practitioners can unleash the power of data science only when we make it safe and find a way to fit it into normal business processes.</p>
5</blockquote>
6<p class="text-gray-600 leading-relaxed mb-4">Here at Kinetica we couldn't agree more! What's the point of building a brilliant model if you can't actually get it into production? That's the problem we're helping to solve by streamlining this machine learning (ML) operationalization process. Kinetica's GPU-accelerated engine can easily and quickly run deep learning models in parallel across multiple machines. Our <a href="https://www.kinetica.com/docs/udf/index.html#udf-label" class="text-kinetica-600 hover:underline">CUDA-optimized user-defined functions API</a>, makes it simple to prepare, train and deploy predictive analytics and ML.</p>
7<p class="text-gray-600 leading-relaxed mb-4">In this blog, we'll go through the process of registering and running a trained image classification model in Kinetica. This is an example of operationalizing ML workloads with a GPU data engine to create a quicker path to value for work coming from an organization's data science teams.</p>
8<p class="text-gray-600 leading-relaxed mb-4"><strong><em><a href="/kinetica-developer-edition" class="text-kinetica-600 hover:underline">Download a trial of the Kinetica engine.</a></em></strong></p>
9
10<h2 class="text-2xl font-bold text-gray-900 mt-8 mb-4">Model Details</h2>
11<p class="text-gray-600 leading-relaxed mb-4">The model used here is trained through transfer learning, the base model is <a href="https://www.tensorflow.org/hub/modules/image" class="text-kinetica-600 hover:underline">Mobilenet</a>. The training dataset is <a href="https://www.vision.ee.ethz.ch/datasets_extra/food-101/" class="text-kinetica-600 hover:underline">Food101</a>. The model has a top-1 accuracy around 71%. In this example, the model itself is a standalone Google <a href="https://developers.google.com/protocol-buffers/docs/overview" class="text-kinetica-600 hover:underline">protobuf</a> file. For TensorFlow, the advantage of using protobuf is that it contains both graph definition and parameter weights. In other words, you can use a single protobuf file to load the whole model graph.</p>
12
13<h2 class="text-2xl font-bold text-gray-900 mt-8 mb-4">Preparing the Input and Output Tables</h2>
14<p class="text-gray-600 leading-relaxed mb-4">To prepare an input table, we can use the <a href="https://www.kinetica.com/docs/tools/kifs.html?highlight=kifs" class="text-kinetica-600 hover:underline">Kinetica File System (KiFS)</a> to create a table named <strong>foodimages</strong> and upload images through the UI. You can check out my colleague's previous <a href="/blog/handling-image-data-kinetica-file-system-kifs" class="text-kinetica-600 hover:underline">blog</a> for details about using KiFS.</p>
15<p class="text-gray-600 leading-relaxed mb-4">For the output table, we'll use the following SQL to create an empty table:</p>
16<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im1.png" alt="" width="375" height="78" class="rounded-lg shadow-md my-6 max-w-full h-auto">
17
18<h2 class="text-2xl font-bold text-gray-900 mt-8 mb-4">Register the User-Defined Function (UDF)</h2>
19<p class="text-gray-600 leading-relaxed mb-4">Before registering the UDF, make sure: cluster→ conf→ &nbsp;enable_procs = true, if you're using a GPU instance, read the "Run on GPU" instructions for setup.</p>
20<p class="text-gray-600 leading-relaxed mb-4">You can use the UI to register UDF as following:</p>
21<p class="text-gray-600 leading-relaxed mb-4">
22  <code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">Name: foodrec</code><br>
23  <code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">Command: python</code><br>
24  <code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">Arguments: Food_predict_kifs.py</code><br>
25  <code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">Distributed: yes</code><br>
26  <code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">Files: food101.pb, food101.txt, Food_predict_kifs.py</code>
27</p>
28<p class="text-gray-600 leading-relaxed mb-4"><a href="https://drive.google.com/open?id=19YhfNM7jK5G1Wq0lJNswhlAuxqo4LnaG" class="text-kinetica-600 hover:underline">Download code and model files</a></p>
29<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/Screen-Shot-2018-08-30-at-2.31.37-PM-1024x620.png" alt="" width="1024" height="620" class="rounded-lg shadow-md my-6 max-w-full h-auto">
30
31<h2 class="text-2xl font-bold text-gray-900 mt-8 mb-4">Execute UDF</h2>
32<p class="text-gray-600 leading-relaxed mb-4">Using the UI, go to UDF→ UDF. Select foodrec, click the execute button, specify input table = foodimages, output table = foodclassification, then click the execute button in the small window.</p>
33<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im3.png" alt="" width="394" height="584" class="rounded-lg shadow-md my-6 max-w-full h-auto">
34
35<h2 class="text-2xl font-bold text-gray-900 mt-8 mb-4">UDF details</h2>
36<p class="text-gray-600 leading-relaxed mb-4">
36Now let's take a look at the UDF itself. There are three files we registered:</p>
37<p class="text-gray-600 leading-relaxed mb-4">
38  <strong>Food101.pb:</strong> Tensorflow model in protobuf format.<br>
39  <strong>Food101.txt:</strong> Contains all the class names, in this example, food names.<br>
40  <strong>Food_predict_kifs.py:</strong> UDF code.
41</p>
42<p class="text-gray-600 leading-relaxed mb-4">For the UDF code, function <strong>readLabel</strong> is used to read the <strong>Food101.txt</strong> file into a python dictionary object like <em>{0:"apple pie", 1:"baby back ribs"…}.</em></p>
43<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im4.png" alt="" width="468" height="209" class="rounded-lg shadow-md my-6 max-w-full h-auto">
44<p class="text-gray-600 leading-relaxed mb-4">The function <strong>load_graph</strong> loads the protobuf object into a tensorflow default graph. The input tensor name is <strong>'Placeholder:0'</strong> with shape . This allows a batch of images with width=224, height=224, channel number=3. The output tensor name is <strong>'final_result:0'</strong> with shape .</p>
45<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im5.png" alt="" width="590" height="135" class="rounded-lg shadow-md my-6 max-w-full h-auto">
46<p class="text-gray-600 leading-relaxed mb-4">As the general images may have arbitrary width and height instead of 224×224, we need to get a tensorflow subgraph to do the transform. The function <strong>transform</strong> builds a subgraph which takes in an image in binary bytes with any width and height and transforms it to 224×224.</p>
47<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im6.png" alt="" width="520" height="91" class="rounded-lg shadow-md my-6 max-w-full h-auto">
48<p class="text-gray-600 leading-relaxed mb-4">The function <strong>run_inference_on_image</strong> takes in an image in bytes and outputs the class index number. For example, if the prediction is <em>baby back ribs</em>, the returned number will be 1.</p>
49<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im7.png" alt="" width="561" height="300" class="rounded-lg shadow-md my-6 max-w-full h-auto">
50<p class="text-gray-600 leading-relaxed mb-4">All the four functions above are standalone modules. It can be used outside Kinetica for testing. The <strong>main</strong> function below contains the pipeline to load data from Kinetica database and write it back to the database. In every UDF, <strong>proc_data = ProcData()</strong> needs to be called before loading input data and <strong>proc_data.complete()</strong> needs to be called after writing output data into output table. Note here we use a for loop to go through each image. This design is to make sure it can run from a very small Kinetica instance (for example, the CPU version on a laptop) to a large cluster without incurring an "out of memory" error. It's very straightforward to switch to batch mode instead of using the for loop.</p>
51<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im8.png" alt="" width="614" height="271" class="rounded-lg shadow-md my-6 max-w-full h-auto">
52
53<h2 class="text-2xl font-bold text-gray-900 mt-8 mb-4">Run on GPU</h2>
54<p class="text-gray-600 leading-relaxed mb-4">This UDF runs on CPU by default so if you're running Kinetica on GPUs like most of our users, you'll need to do some extra steps to get it set up.</p>
55<p class="text-gray-600 leading-relaxed mb-4"><strong>Shutdown Kinetica, modify Cluster→ conf:</strong></p>
56<p class="text-gray-600 leading-relaxed mb-4">
57  enable_procs = true<br>
58  enable_gpu_allocator = false
59</p>
60<p class="text-gray-600 leading-relaxed mb-4"><strong>Modify Food_predict_kifs.py:</strong><br>
61By default, this UDF only runs on CPU, from line 7:<br>
62<code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">os.environ["CUDA_VISIBLE_DEVICES"]=""</code></p>
63<p class="text-gray-600 leading-relaxed mb-4">To modify it to use certain GPUs (for example, GPU with index 0):<br>
64<code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">os.environ["CUDA_VISIBLE_DEVICES"]="0"</code><br>
65If your cluster has multiple nodes, and each node has multiple GPU cards, this will force every UDF to use its number 0 GPU card on its node.</p>
66<p class="text-gray-600 leading-relaxed mb-4">If all the GPU cards are available, you can set it as following:</p>
67<p class="text-gray-600 leading-relaxed mb-4"><code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">os.environ['CUDA_VISIBLE_DEVICES'] = str(int(proc_data.request_info["rank_number"]) – 1)</code></p>
68<p class="text-gray-600 leading-relaxed mb-4">This way each TOM will use the GPU that's attached to its own rank.</p>
69<p class="text-gray-600 leading-relaxed mb-4">Another thing to note is: by default TensorFlow will try to grab all the GPU for each session, this can easily cause an CUDA "out of memory" error. To avoid it we need the following:</p>
70<img loading="lazy" decoding="async" src="https://kinetica-web-assets.s3.us-east-1.amazonaws.com/assets/blog/im9.png" alt="" width="347" height="35" class="rounded-lg shadow-md my-6 max-w-full h-auto">
71<p class="text-gray-600 leading-relaxed mb-4">This won't be effective by default unless you change the following in the main function:<br>
72<code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">with tf.Session() as sess:</code> &nbsp;→ <code class="bg-gray-100 text-kinetica-700 rounded px-1.5 py-0.5 text-sm font-mono">with tf.Session(config=gpu_options) as sess:</code></p>
73
74<h2 class="text-2xl font-bold text-gray-900 mt-8 mb-4">Conclusion</h2>
75<p class="text-gray-600 leading-relaxed mb-4">I hope this blog provided a helpful guide on registering your own model in Kinetica! The ability to run your AI and ML workloads on Kinetica alongside streaming and geospatial analytics is incredibly powerful. Keeping all of your data operations on a single engine increases time to insight and enables a wider range of users to leverage the power of your data, from business analysts to data scientists. If you have any problems or questions please leave a comment.</p>
76<p class="text-gray-600 leading-relaxed mb-4"><strong>Additional online resources are available here:</strong></p>
77<ul class="list-disc pl-6 mb-4 space-y-2 text-gray-600">
78  <li><a href="/kinetica-developer-edition" class="text-kinetica-600 hover:underline">Get the Kinetica trial license key</a></li>
79  <li><a href="https://hub.docker.com/r/rewreu/kinetica-intel/" class="text-kinetica-600 hover:underline">Kinetica Docker images for CPU version</a></li>
80  <li><a href="https://hub.docker.com/r/kinetica/kinetica-cuda91/" class="text-kinetica-600 hover:underline">Kinetica Docker images for GPU version</a></li>
81</ul>
82<p class="text-gray-600 leading-relaxed mb-4"><em><strong>Zhe Wu is a Data Scientist at Kinetica.</strong></em></p>`,l={slug:e,title:t,excerpt:a,date:n,author:s,featuredImage:o,categories:i,body:r};export{s as author,r as body,i as categories,n as date,l as default,a as excerpt,o as featuredImage,e as slug,t as title};

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.