PageSourceSearch

https://opentsdb.net/faq.html

html opentsdb.net collected 2026-09-24 08:28:18 UTC 21,188 bytes, 381 lines download raw bytes

1<!DOCTYPE html>
2<html lang="en">
3  <head>
4    <meta charset="utf-8">
5    <meta http-equiv="X-UA-Compatible" content="IE=edge">
6    <meta name="viewport" content="width=device-width, initial-scale=1">
7    
8    <title>FAQ - OpenTSDB - A Distributed, Scalable Monitoring System</title>
9    
10    <!-- Bootstrap core CSS -->
11    <link href="css/bootstrap.min.css" rel="stylesheet">
12    <link href="css/otsdb-bootstrap.css" rel="stylesheet">
13
14    <!-- Custom styles for this template -->
15    <link href="css/custom.css" rel="stylesheet">
16
17    
17<script type="text/javascript">
18      var _gaq = _gaq || [];
19      _gaq.push(['_setAccount', 'UA-18339382-1']);
20      _gaq.push(['_setDomainName', 'none']);
21      _gaq.push(['_setAllowLinker', true]);
22      _gaq.push(['_trackPageview']);
23      (function() {
24        var ga = document.createElement('script'); ga.type = 'text/javascript'; ga.async = true;
25        ga.src = ('https:' == document.location.protocol ? 'https://ssl' : 'http://www') + '.google-analytics.com/ga.js';
26        var s = document.getElementsByTagName('script')[0]; s.parentNode.insertBefore(ga, s);
27      })();
28    </script>
28
29  </head>
30
31  <body>
32     <div class="container">
33       <h1><a href="http://www.opentsdb.net/"><img src="img/logo_header.png"></a></h1>
34       
35     <div class="navbar navbar-default">
36        <div class="container-fluid">
37          <div class="navbar-header">
38            <ul class="nav navbar-nav">
39              <li><a href="index.html">Home</a></li>
40              <li><a href="overview.html">Overview</a></li>
41              <li><a href="https://github.com/OpenTSDB/opentsdb/releases">Download</a></li>
42              <li><a href="https://github.com/OpenTSDB/opentsdb">Source</a></li>
43              <li><a href="docs/build/html/index.html">Documentation</a></li>
44              <li><a href="faq.html">FAQ</a></li>
45            </ul>
46          </div>
47        </div>
48      </div>
48<script src="misc/toc.js" type="text/javascript"></script>
48
49<section id="FAQ">
50<div id="toc"></div>
51
52<a name="scalability"></a><h2>Scalability</h2>
53<h4>Can OpenTSDB scale to multiple data centers?</h4>
54Yes.  It is recommended that you run one set of Time Series Daemons
55(TSDs) per HBase cluster and one HBase cluster per physical datacenter.
56It is not recommended to have HBase clusters spanning across data
57centers.  Instead you can use
58<a href="http://hbase.apache.org/replication.html">HBase replication</a>
59to replicate tables across data centers.
60
61<h4>How much <em>write</em> throughput can I get with OpenTSDB?</h4>
62It depends mostly on two things:
63<ol><li>The size of your HBase cluster.</li><li>The CPUs you're using.</li></ol>
64If your HBase cluster is reasonably sized, it's unlikely that OpenTSDB will
65max it out as the TSDs tend to be CPU bound before that happens (unless you
66run many TSDs).  A TSD can easily handle 2000 new data points per second per
67core on an old dual-core Intel Xeon CPU from 2006.  More modern CPUs will get
68you more throughput.
69
70<h4>How much <em>read</em> throughput can I get with OpenTSDB?</h4>
71Read throughput varies depending on the cardinality of a metric (how many 
72distinct time series exist), the time span and the number of data points retreived. 
73Lower cardinality with fewer data points will execute much quicker than higher
74cardinality and greater data point queries. Most queries for the last day of data
75will return in less than a second with low cardinality. However huge queries can
76run for multiple seconds. We're working to optimize the query path.
77
78<h4>What type of hardware should I run the TSDs on?</h4>
79There are no strict requirements.  The recommended configuration, however, is
80a 4-core machine with at least 4GB of RAM, and a <code>tmpfs</code> partition
81for the cache directory used by the TSD.  Having more RAM helps the TSD ride
82over transient HBase outages by allowing it to buffer more incoming data
83before getting to the point where it must start discarding data.
84<p>
85HBase region servers are usually beefy machines and many OpenTSDB users run
86their TSDs on the same machines, granted there is enough memory for both processes.
87
88<h4>How much disk space do I need?</h4>
89The answer depends mostly on the average number of tags per data point. 
90StumbleUpon uses 4.5 tags on average and 100+ billion data points take
91only just over a terabyte of disk space (pre-HDFS 3x replication).  Enabling compression 
92with <a href="setup-hbase.html#lzo">LZO</a> or Snappy is
93extremely recommended in a production setting.  In Stumbleupon's case, each data point
94ends up taking about 12 bytes of disk space (or actually 36 if you include the
953x replication factor of HDFS).  We also find that, on average, LZO is able to
96achieve a compression factor of 4.2x on the TSD table, but your mileage will vary.
97Without LZO, a data point costs roughly: 16 bytes of HBase overhead, 3 bytes
98for the metric, 4 bytes for the timestamp, 6 bytes per tag, 2 bytes of
99OpenTSDB overhead, up to 8 bytes for the value. Integers are stored with variable
100length encoding and can consume 1, 2, 4 or 8 bytes.
101
102<a name="reliability"></a><h2>Reliability</h2>
103<h4>What are the Single Points of Failure of OpenTSDB?</h4>
104OpenTSDB itself doesn't have any specific SPoF as there is no central
105component and you can run multiple TSDs on different machines.  The TSDs
106need HBase to run, and HBase doesn't have any SPoF<sup>*</sup> either as HBase
107only really needs a ZooKeeper quorum to keep serving.  A ZooKeeper quorum is
108typically made of 5 different machines, out of which you can afford to lose 2
109before the system goes down.  Note that although HBase has a master, it is not
110actually needed for HBase to keep serving.  Not having a master running will
111prevent HBase from starting or recovering from machine failures but, in a
112steady state, losing the master doesn't impede on HBase's ability to serve.
113<p>
114<small><sup>*</sup> Fine prints: if your HBase cluster is backed by HDFS,
115which is most likely the case for production clusters at the time of this
116writing, then you have a SPoF because of the NameNode of HDFS.  If you run
117HBase on top of a <em>reliable</em> distributed filesystem, then you don't
118have any SPoF.</small>
119
120<h4>What are the failure modes of OpenTSDB?</h4>
121The TSD eventually becomes unhealthy when HBase itself is down or broken for
122an extended period of time.  Right now, the TSD doesn't handle 
123prolonged HBase outages very well and will discard incoming data points once its buffers
124are full if it's unable to flush them to HBase.  Future versions will temporarily store
125data to local disk when it is unable to reach HBase.
126<p>
127StumbleUpon has had a number of cases where a collector that runs on
128hundreds of machines goes crazy and generate a DDoS on the TSDs.  The TSD
129doesn't do a good job at handling such DDoS situations by penalizing offending
130clients, so its performance will degrade once the machine it's running on is
131unable to keep up with the load.
132
133<h4>What is the recommended deployment for OpenTSDB?</h4>
134We recommend that you run multiple TSDs behind a load balancer such as Varnish, 
135HAProxy or DNS round robin. StumbleUpon found it useful
136to dedicate one or more TSDs for read queries (human users using the web UI to
137generate graphs or to view dashboards) and let other TSDs handle the write
138queries (new data points coming in from production machines).  For the
139"read-only" TSDs, we recommend Varnish for load balancing.
140<a href="varnish.html">Read more about Varnish and TSDs</a>.
141
142<h4>What data durability guarantees does OpenTSDB make?</h4>
143By default the TSD buffers data points for about 1 second before persisting
144them in HBase (configurable via the <code>--flush-interval</code> flag).  If
145the TSD was to crash without getting a chance to run its shutdown hook, you
146could lose up to 1 second worth of data points.  In practice we've found this
147trade off to be acceptable given the performance benefits that deferred
148flushes offer in terms of write throughput.  Once a data point has been stored
149in HBase, data durability is guaranteed if you're running HBase on top of a
150distributed filesystem that provides the necessary data durability guarantees.
151<p>
152If you use HDFS, we recommend that you run
153<a href="http://www.cloudera.com/downloads/">Cloudera's Distribution for
154Hadoop</a> (CDH), version 3 or above preferably, as this version comes with
155all the necessary patches to make HDFS less unreliable and has better
156performance.
157
158<a name="datamodel"></a><h2>Data Model</h2>
159<h4>How can I increment a counter in OpenTSDB?</h4>
160OpenTSDB does not have a counter feature at this time, though work is underway.  
161Currently OpenTSDB simply records <code>(timestamp,
162value)</code> pairs.  Data points are independent from each other.  Say you
163want to keep track of clicks on an ad in OpenTSDB.  You wouldn't send a "+1"
164to the TSD for every click.  Instead, if your application doesn't already
165keeps track of click counts, you'd need to increment a counter for every click
166and periodically send the value of that counter to the TSD.  You can later
167query the TSD and ask for the rate of change of the counter, which will give
168you clicks per second.
169
170<h4>Can I store sub-second precision timestamps in OpenTSDB?</h4>
171As of version 2.0 you can store data with millisecond timestamps. However 
172we recommend you avoid storing a data point at each millisecond as this will
173slow down queries dramatically. See 
174<a href="docs/build/html/user_guide/query/dates.html">Dates and Times</a> 
175for details.
176
177<h4>Can I use another storage backend than HBase?</h4>
178Not at this time.  OpenTSDB was designed specifically for a storage backend that follows
179the <a href="http://labs.google.com/papers/bigtable.html">Bigtable</a> data
180model (a distributed, strongly consistent, sorted multi-dimensional hash map).
181At the time OpenTSDB was written, HBase is the only such system that's both open-source
182and usable in production, so the code was written specifically for HBase.
183Technically it would be feasible to port the code to other systems that follow
184the Bigtable data model.  Systems that differ by not storing data in a sorted
185fashion (such as distributed hash tables) or that do not offer a strong
186consistency guarantee will simply not work with the current design.
187<p>
188We're looking at adding support for Cassandra now that it implements counters.
189
190<a name="misc"></a><h2>Misc</h2>
191<h4>How do the TSDs handle DST changes or leap seconds?</h4>
192The TSD doesn't assign timestamps to your data points, your collectors do.
193It is strongly recommended that you use UNIX timestamps in your collectors,
194so all your timestamps will be based on
195<a href="http://en.wikipedia.org/wiki/Unix_epoch">Epoch</a>.  This way
196you will not be affected by timezone adjustments or DST changes on your
197machines.
198<p>
199The TSD always renders timestamps in local time when using the GUI, to make it easier for us
200human to understand and correlate events based on the timezone we live in.
201So you should to make sure you give the TSD the correct timezone setting
202(e.g. via the <a href="cli.html#Overriding_the_timezone_of_the_TSD"><code>TZ</code> environment variable</a>).
203When the TSD starts,
204it computes its offset from UTC and will then keep that offset forever.
205In case of a DST change, for instance, it would then appear that the TSD
206is 1 hour behind.  There are plans to periodically re-compute the offset
207from UTC to avoid that situation, but right now you have to restart the
208TSD in order to adjust the offset.  Note that this doesn't prevent the TSD
209from working properly, it only affects anything that parses dates from local
210time or renders them in local time.  Dashboards and alerting systems should
211use relative time (e.g. "1d ago") and should thus be unaffected.
212<p>
213When <a href="http://en.wikipedia.org/wiki/Leap_second">leap</a> seconds
214occur, UNIX timestamps go back by one second.  The TSD should handle this
215situation gracefully (although this hasn't been tested yet).  Unless you're
216collecting data every second, you won't notice anything except that the
217interval between the two data points where the leap second occurred is one
218second less than it should have been.  If you do collect data every second,
219the second data point that attempts to overwrite the previous one during the
220leap second will be discarded with an error message.
221
222<h4>The graphs are ugly, can they be made prettier?</h4>
223Ugliness is a subjective thing :)
224<p>
225There are a lot of knobs that aren't exposed yet that would allow the TSD to
226generate nicer, antialiased, smoothed graphs.  It's just a matter of exposing
227those Gnuplot knobs.  Also, recent versions of Gnuplot can generate graphs in
228HTML5 canvas.  We plan to use this to build pretty graphs you can interact
229with from your web browser.
230<p>
231Please contribute to help make the UI sexier.
232
233<h4>Can I use OpenTSDB to generate graphs for my customers / end-users?</h4>
234Yes, but you have to be careful with that.  OpenTSDB was written for internal
235use only, to help engineers and operations staff understand and manage large
236computer systems.  It hasn't been through any security review and does not 
237included authentication.
238<p>
239We don't recommend that you give direct access to the TSD to untrusted users.
240If you really want to leverage the TSD's graphing features, we recommend that
241you put the TSD behind a secured HTTP proxy that only allows specific requests
242to go through.  Alternatively, you could use the TSD to periodically
243pre-generate a fixed set of graphs and serve them as static images to your
244customers.
245
246<h4>Why does OpenTSDB return more data than I asked for in my query?</h4>
247All queries specify a start time and an end time (if the end time isn't
248specified, it is assumed to be "now").  OpenTSDB's goal is to plot a sensible
249graph covering that time span.  However it needs to retrieve data before and
250after the times you actually specified, in order to know how to properly
251compute the values near the "edges" of the graph.  Because having extra values
252past the times actually requested is required to draw accurate graphs,
253OpenTSDB also returns the extra data based on the assumption that if you want
254to plot your own graphs or make your own processing, you will also need the
255extra data to get the correct behavior near the edges.
256<p>
257The amount of extra data that OpenTSDB attempts to retrieve is proportional to
258the time span covered by your query.  The 2.0 HTTP API will only return data within 
259the requested time span.
260
261<h4>I don't understand the data points returned for my query</h4>
262Sometimes the results to a query don't match people's expectations.  This is
263often because it's not necesssarily quite obvious what steps are involved in
264a query, why OpenTSDB uses interpolation, when do aggregators kick in.
265Please see the documentation for 
266<a href="http://opentsdb.net/docs/build/html/user_guide/query/aggregators.html">Aggregators</a> 
267for details.
268
269<a name="meta"></a><h2>Meta</h2>
270
271<h4>Why was OpenTSDB written in Java?</h4>
272Mostly because OpenTSDB lives around the HBase community, which is a Java
273community.  OpenTSDB also started by using HBase's library, which only
274exists in Java.  Eventually, however, OpenTSDB started to use another
275alternative library to access HBase
276(<a href="http://github.com/OpenTSDB/asynchbase">asynchbase</a>)
277but sadly this one too is in Java.
278
279<h4>Why HBase and not Cassandra</h4>
280Early on Cassandra didn't support atomic operations required in the assignment of UIDs.  
281Now that Cassandra has such features, including counters, it would be possible to copy
282the HBase schema in Cassandra, though there are some differences to work through and
283an asynchronous driver may need to be developed.
284
285<h4>Has OpenTSDB anything to do with OpenBSD?</h4>
286While the author of OpenTSDB admires the work done on OpenBSD, the fact that
287the name of projects are so close is just a coincidence.  "TSDB" alone was
288too ambiguous, and the author miserably failed to come up with a better name.
289
290<h4>Who is behind OpenTSDB?</h4>
291OpenTSDB was originally designed and implemented at
292<a href="http://www.stumbleupon.com">StumbleUpon</a> by Beno&icirc;t Sigoure.
293Berk D. Demir contributed ideas and feedback during the early design stages.
294<a href="https://github.com/OpenTSDB/tcollector"><code>tcollector</code></a>
295was designed and implemented by Mark Smith and Dave
296Barr, with contributions from Beno&icirc;t Sigoure.
297
298<h4>How to contribute to OpenTSDB?</h4>
299The easiest way is to <a href="http://help.github.com/fork-a-repo/">fork</a>
300the project on <a href="https://github.com/OpenTSDB/opentsdb">GitHub</a>.
301Make whatever changes you want to your own fork, then send a
302<a href="http://help.github.com/send-pull-requests/">pull request</a>.
303You can also send your patches to the
304<a href="http://groups.google.com/group/opentsdb">mailing list</a>.
305Be prepared to go through a couple iterations as the code is being reviewed
306before getting accepted in the main repository.  If you are familiar with
307how the Linux kernel is developped, then this is pretty similar.
308
309<h4>Who commits to OpenTSDB?</h4>
310Anyone can commit to OpenTSDB, provided that the changes are accepted
311in the main repository after getting reviewed.  There is no notion of
312"commit access", or no list of committers.
313
314<h4>Why does OpenTSDB use the LGPL?</h4>
315One of the most frequent "holy war" that plague open-source communities is
316that of what licenses to use, which ones are better or "more free" than others.
317OpenTSDB uses the <a href="http://www.gnu.org/licenses/lgpl.html">GNU LGPLv2.1+</a>
318for maximum compatibility with its dependencies and other licenses, and
319because its author thinks that the LGPL strikes the right balance between the
320goals of free software and the legal restrictions often present in corporate
321environments.
322<p>
323Let's stress the following points:
324<ul>
325<li>The LGPL is <em>not</em> the GPL.  Although based on the same text, the
326way it extends the GPL has significant consequences.  Do not confuse the two.</li>
327<li>The <a href="http://www.gnu.org/licenses/lgpl-java.html">LGPL is perfectly
328compatible with Java</a>.  The myth that the LGPL does not work as intended
329with Java is, well, just a myth, albeit a widespread one.</li>
330<li>The LGPL allows you to use the code in proprietary software, provided that
331you don't <em>redistribute a modified version</em> of the LGPL'ed code.</li>
332<li>If you want to redistribute a modified version of the code, then your
333changes must be released under the LGPL.</li>
334<li>The LGPL is perfectly compatible with the
335<a href="http://www.apache.org/licenses/LICENSE-2.0">ASF2</a> license.
336Many people are misled to believe that there is an incompatibility because the
337Apache Software Foundation (ASF) decided to not allow inclusion of LGPL'ed
338code in its own projects.  This choice only applies to the projects managed by
339the ASF itself and doesn't stem from any license incompability.</li>
340</ul>
341With this out of the way, we hope that those afraid of the 3 letters "GPL"
342will acknowledge the importance of using the LGPL in OpenTSDB and will
343overcome their fear of the license.
344<br/>
345<small>Disclaimer: This page doesn't provide any formal legal advice.
346Information given here is given in good faith.  If you have any doubt, talk to
347a lawyer first.  In the text above "LGPL" or "GPL" refers to the version 2.1 of
348the license, or (at your option) any later version.</small>
349
350<h4>Who supports OpenTSDB?</h4>
351<a href="http://www.stumbleupon.jobs">StumbleUpon</a> supported the initial
352development of OpenTSDB as well as its open-source release.
353<a href="http://www.yahoo.com">Yahoo!</a> maintained OpenTSDB until 2021. It is
354currently in maintenance mode as the number of time series storage solutions have
355exploded.
356Many other engineers and companies have contributed to OpenTSDB over the years.
357Please see our <a href="https://github.com/OpenTSDB/opentsdb/blob/next/THANKS">
358thanks</a> file for contributors. And thank you to everyone who has made OpenTSDB
359so popular and successful.
360<p>
361YourKit is kindly supporting open source projects with its full-featured Java
362Profiler.  YourKit, LLC is the creator of innovative and intelligent tools for
363profiling Java and .NET applications. Take a look at YourKit's leading
364software products:
365<a href="http://www.yourkit.com/java/profiler/index.jsp">YourKit Java Profiler</a>
366and
367<a href="http://www.yourkit.com/.net/profiler/index.jsp">YourKit .NET Profiler</a>.
368
369
370</section>
371      <hr>
372      <div class="footer">
373        <p><a href="https://groups.google.com/forum/#!forum/opentsdb"><i class="glyphicon glyphicon-envelope"></i> Mailing List</a>  | <i class="glyphicon glyphicon-user"></i> IRC: Freenode #opentsdb  |  &copy; 2010 - 
373<script type="text/JavaScript"> document.write(new Date().getFullYear()); </script>
373 The OpenTSDB Authors | Travis CI Status: <img src="https://travis-ci.org/OpenTSDB/opentsdb.svg?branch=master">
374      </div>
375    </div> <!-- /container -->
376    <!-- jQuery (necessary for Bootstrap's JavaScript plugins) -->
377    
377<script src="https://ajax.googleapis.com/ajax/libs/jquery/1.11.0/jquery.min.js"></script>
377
378    <!-- Include all compiled plugins (below), or include individual files as needed -->
379    
379<script src="js/bootstrap.min.js"></script>
379
380  </body>
381</html>

Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.