When an application is split into many services, teams often need a consistent way to encrypt the connections between them, decide which service may call another and see where requests fail. A service mesh adds networking components that provide these features alongside the applications, reducing how much each application has to implement itself.

Istio provides traffic routing, identity and security controls using either proxies beside applications or its ambient approach. Linkerd also uses proxies and a management service to provide features such as encrypted service-to-service connections and traffic measurements. In Istio versus Linkerd, compare the features your services need and the extra software your team can maintain. A mesh is unnecessary if ordinary Kubernetes networking, application libraries and your gateway already solve the problem.

Compare deployment models

Istio offers sidecar and ambient approaches; ambient reached general availability in Istio 1.24. Linkerd uses its own proxy/control-plane model. Do not infer feature parity or resource savings from architecture labels. Check the exact release, supported features and distribution channel for each candidate.

  • Identity and encryption: define trust domains, certificate rotation, outside-cluster traffic and the behavior of unmeshed clients.
  • Policy: distinguish connection-level policy from HTTP-aware authorization and routing.
  • Operations: identify who upgrades proxies, control planes, CRDs and workload configuration.
  • Failure: determine how new Pods and established traffic behave when certificate or control-plane services are unavailable.

A useful evaluation

Use the same small application and traffic mix. Check mTLS identity, authorization failures, telemetry attribution and rollout behavior. Measure CPU, memory and tail latency on declared hardware; do not compare a vendor's best-case benchmark with an unrelated default deployment. Include protocol exceptions, health checks and traffic to external services.

Start small and retain an exit

Introduce the mesh into one namespace or service boundary, observe it, then expand. Removal may require changing traffic policies and certificate assumptions as well as deleting proxies. Keep ordinary application connectivity working during a rollback rehearsal.

The legacy URL names Linkerd 1, Linkerd 2 and Consul because it preserves a historical comparison. Linkerd 1's old repository points readers to Linkerd 2; they are not interchangeable generations. Current Consul deployment and support claims require their own version-specific review. This page supplies evaluation criteria, not an unperformed mesh benchmark.

Which should you evaluate first?

Evaluate Istio when you want to combine detailed request routing and authorization with a choice of sidecar or ambient deployment. For example, a staged rollout can send selected traffic to a new version while policy restricts which services may call it. In ambient mode, HTTP-aware routing and policy need the appropriate waypoint proxies. Evaluate Linkerd when the immediate goal is encrypted connections between services and request metrics, and its sidecar approach is acceptable. Its feature documentation also includes routing and authorization: do not assume those capabilities are exclusive to Istio. Choose neither until you can name a problem the mesh will solve. Test that problem with one service pair before accepting the cost of operating another shared system.

The original record

Historical Kubedex content

Original publication: 2018-09-10T13:03:41+00:00. Preserved for context. Commands, versions, prices and results below reflect the original research.

Reading Time: 5 minutes

Last Updated on August 2, 2021

We’re going to compare every Kubernetes service mesh available today and work out who the winner is. You may have already read our Top10 list of Kubernetes applications in which case the result may be somewhat predictable.

If you’ve arrived on this page you probably already understand what a service mesh does. If you don’t then go and quickly read this article and then come back.

In this blog we’ll hopefully help you to choose from the four options available today. Like all of the content on this site it will be 100% Kubernetes specific.

To help inform people about what service mesh to choose I’ve put together a quick table of features.

Historical illustration

Many of these are so new that the documentation is lagging a little behind. If you can help fill in the gaps please add a comment to this spreadsheet and I’ll update this blog accordingly.

Linkerd

I used Linkerd extensively on DC/OS and absolutely loved it. However, times have changed and there are a couple of fundamental problems that have caused this to be a total dead-end on Kubernetes.

Linkerd is written in a JVM language which means a footprint of 110mb+ memory usage per node agent. This isn’t too bad when you just run one node agent per host, but the world is moving to per pod proxy sidecars, and I think everyone realised this is too much overhead.

Linkerd also doesn’t proxy TCP requests and doesn’t support websockets.

On the positive end of the scale Linkerd has absolutely amazing traffic control. Read some of the documentation around Namerd and you’ll see just how advanced and powerful it is. It’s also one of the two service meshes that supports connections outside of the cluster.

So in summary I’d say if you only have Kubernetes to worry about then give Linkerd a miss. If you have Linkerd already in other areas and need to connect services on your Kubernetes cluster to them then it may be a valid option.

Linkerd2 (formerly Conduit)

Linkerd2 is a total rewrite of Linkerd in Golang and Rust specifically for Kubernetes. Unfortunately, as with every rewrite, you start back at the beginning again from a feature and stability perspective. Although I’m sure there are more than a few lessons learned.

Moving to Rust for the data plane proxy sidecars should help mitigate some of the bugs and should also solve the memory issues. It also supports all of the major protocols now which is a big step forward.

One interesting difference compared to other service mesh designs is the tight default coupling between the data plane and control plane services. This simplifies the configuration which I see as a positive. I also like how there is a focus on keeping the data plane latency P99’s extremely low.

As mentioned at the beginning though the project just isn’t at a level from a feature perspective where it can compete with something like Istio. To give just one example of something I’d consider to be fundamentally required from a service mesh: distributed tracing.  This is still in in the planning stage for Linkerd2. There are many other features that other blogs have called ‘table stakes’ that seem to be still in the RFC stage.

Having said this, if you try Linkerd2 and are happy with the current feature set then this seems like a good investment for the future. Many people hate the high complexity of Istio and so I think over time this may become the most compelling option if it remains simple.

Update: Recent updates include sidecar injection, timeouts and retries.

Consul

The latest version of Consul now comes with the ‘connect‘ feature which can be enabled on existing clusters. Like with most of the Hashicorp tools Consul is a single Go binary that includes both the data and control plane. The main unique selling point seems to be that you can enable connect across services on Kubernetes and join them to services on vm’s outside that also run Consul. This might be attractive to some organisations. However, I don’t really see it as a big advantage based on work I’ve done in the past. Usually we leave the legacy alone and let it die, then work on new projects or migrate stuff onto Kubernetes.

Consul does seem to have a slight architectural advantage in that it operates as a full mesh with no centralised control plane services that could theoretically act as a performance bottleneck.

There is also a neat separation between layer4 and layer7. I think this separation may keep the Consul service mesh design simple while still allowing for the data layer to be split out. You can currently switch out the default data layer proxy with Envoy if you need more layer 7 features.

The default proxy is however quite lacking in features. To get tracing support, or many of the more advanced layer 7 features, you’ll need to swap out the proxy for something like Envoy. This isn’t very well documented online.

The other part to keep an eye on is the control plane configuration. Istio is notoriously complicated to configure at this layer and I see Consul has a simple ‘service access graph’ feature.

Hashicorp have blogged about differentiating in the area of security. Consul ACL’s providing host to host security is a very nice feature. Especially if you want to connect pods from inside Kubernetes outside the cluster in a secure way. The agent caching, especially for auth, apparently makes the communication performance excellent.

So just like Linkerd2 this is another one to watch. Consul connect was only released a few weeks ago and so there really aren’t many howto guides online. If you’re already highly invested in the Hashicorp toolchain then I’d trial this and perhaps learn about how to swap out the default proxy with Envoy.

Istio

Istio is stable and feature rich. At the time of writing Istio has 11.5k Github stars, 244 contributors and is backed by Lyft, Google and IBM. Istio has pioneered many of the ideas currently being emulated by other service meshes.

One such stand-out-feature is the automatic sidecar injection which works amazingly well with Helm charts.

There are of course some negatives which are all to do with modularity, plug-ability and ultimately complexity. You can switch out almost any component of Istio and integrate it with other systems. This all comes at the cost of a steep learning curve and plenty of scope to shoot yourself in the foot.

However, surprisingly, you can get up and running with Istio very quickly if you stick to the defaults. Configuring a test instance with minikube, helm and Istio on your laptop is less than 5 minutes of work. There are also thousands of articles online for how to configure other integrations. This is a stark contrast to the other service meshes.

The winner: Istio?

It’s close but I’d say if you’re starting from scratch on Kubernetes which many people are then Istio is probably the best service mesh right now. The complexity is high, but not massively high when compared to what you have to manage with Kubernetes already. It has the most features and went version 1 and production ready a few months ago. It has also got the backing of Google and a massive community churning out cool blogs and integrations.

Edit: As of 2021 I’ve been using Linkerd2 in production successfully. Each time we have recently evaluated Istio the operational overhead and steep learning curve was a concern to everyone on the team. Linkerd2 is simple, is now production ready, and we have had no problems with it at all.

The end..

Or probably not. Perhaps this comparison was too premature. It’s nice to have a competitive landscape with software and hopefully one day I can revisit this list and crown a new winner.

If anything I’ve written is technically inaccurate please drop me a message below.

To keep updated as new service meshes are released sign up for our news letter and visit our service mesh category.

 

Historical workbook values; blanks mean unknown, not No. Prices, versions, maturity labels and feature claims are not current recommendations.

Kubernetes Service Mesh / Sheet1 (historical)

Historical snapshot

Recovered comparison data. Versions, prices and availability describe the original research, not a current benchmark. Blank or damaged source values are marked unknown.

Choose columns
Visible columns
Kubernetes Service Mesh / Sheet1 (historical)
Feature / source labelIstioLinkerd 1.x (old)Linkerd 2.xConsul ConnectMaesh (Traefik)Kuma
ModelSidecarNode AgentSidecarSidecarNodes ProxySidecar
PlatformKubernetesAnyKubernetesAnyK8sAny
languageGoJVMGo / RustGoGoGo
ProtocolHTTP1.1 / HTTP2 / gRPC / TCP / UDPHTTP1.1 / HTTP2 / gRPCHTTP1.1 / HTTP2 / gRPC / TCP, websocketHTTP1.1 / HTTP2 / gRPC / TCPHttp1.1 / HTTP2 / TCP / gRPC / WebsocketsHTTP1.1 / HTTP2 / gRPC / TCP / UDP
Default Data PlaneEnvoy (supports others)Nativelinkerd-proxy (Rust)Envoy (supports others)TraefikEnvoy
Sidecar InjectionYesNoYesYesNoYes
EncryptionYes, and on by defaultYesYes, and on by defaultYesNot between pods (acceptable trade off)Yes
Traffic Controllabel/content based routing, traffic shiftingDynamic request routing, traffic shifting, per request routingTraffic shiftingstatic upstream, prepared query, http api / dns with native integration, Traffic Splitting, HTTP path-based routingDNSUnknown
Resiliencetimeouts, retries, connection pools, outlier detectiontimeouts, retries, deadlines, circuit breakingRetries, timeouts, connection poolsRetries, timeouts, circuit breakingRetries,timeouts, breaker circuit, rate limiter etcRetries,timeouts, breaker circuit, rate limiter etc
Prometheus IntegrationYesYesYesYesYesYes
Tracing IntegrationJaegerZipkinYesPluggableOpenTracing (Zipkin/Jaeger etc)Unknown
Host to Host authService AccountsTLS Mutual AuthService accounts, mTLSConsul ACLService AccountsUnknown
Agent CachingYesNoYesYes?Unknown
Secure connection outside clusterNoYesYesYesUnknownUnknown
ComplexityHighHighLowLowLowLow
Resource costHighHighLow?LowLow
Latency addedLowHighLow?LowLow
Paid SupportYesYesYesYesYesYes
Hosted versions14.00.00.00.00.01.0
linkhttps://istio.io/https://linkerd.io/1/overview/https://linkerd.io/2/overview/https://www.consul.io/mesh.htmlhttps://docs.mae.sh/install/Unknown

20 rows

Sources & further reading

  1. Istio ambient GA
  2. Linkerd architecture overview
  3. Linkerd 1 historical repository
  4. Recovered historical source (Common Crawl index)
  5. Kubernetes Service Mesh: historical workbook
  6. Istio sidecar and ambient modes
  7. Istio request authorization
  8. Linkerd features

Spotted something that needs another look?

Help improve this page →