Showing posts with label Internet. Show all posts
Showing posts with label Internet. Show all posts

Friday, March 26, 2010

Pastikan anak anda aman saat online

 

os-articleimg-childsafe1

Kebebasan untuk menjelajahi informasi tak terbatas sangat luar biasa dan mengasyikkan bagi anak-anak. Tapi ingat, bahaya menanti di setiap sudut internet. Bantu anak Anda menggunakan internet secara aman dengan memberitahukan beberapa peraturan mendasar.

Coba untuk membuat sebuah tata krama dan tata cara bagi keluarga berupa aturan untuk disepakati bersama.

1. JANGAN pernah berbagi password dan tetap rahasiakan

2. Jangan melakukan kontak online dengan orang tidak dikenal

3. Jatahkan dan pantau penggunaan internet bagi anak-anak Anda

    • Sampaikan ke anak Anda mana yang boleh dan tidak boleh dibuka dan seberapa lama mereka bisa gunakan internet
    • Anda bisa mencegah (meski tidak 100%) akses ke internet dengan menggunakan content filtering gratis dari Blue Coat - K9 atau Windows Live Family Safety

4. Tingkatkan pengaman online Anda

Tentukan batasan untuk downloading dan instal perangkat anti-spyware

5. Gunangan situs jaringan sosial secara aman

    • Hati-hati memasukkan dan berbagi gambar atau informasi secara online
    • Hati-hati atas ekspresi kata-kata anak-anak terhadap orang tak dikenal
    • Berjaga-jagalah terhadap bullying melalui internet
    • Jika Anda merasa jaringan pengaman Anda terancam, hubungi situs web jaringan sosial yang dipakai anak Anda dan ajukan permintaan untuk mencabut halaman tersebut. Juga, cobalah alat Internet-filtering agar bisa memantau aktivitas online anak Anda.

6. Jangan terlalu membuka diri di dalam situs blog

    • Pilah apa yang akan di posting oleh anak Anda. Jangan posting hal-hal yang Anda rasa kurang nyaman
    • Evaluasi layanan blogging dan cari tahu apakah menyediakan layanan blogging dengan pengaman password pribadi
    • Pantau blogging anak Anda secara berkala
    • Cari situs blog yang baik dan bermanfaat bagi anak untuk menambah pengetahuan

7. Hati-hati terhadap penipuan online

    • Jagalah data pribadi yang penting
    • Log-off jika gunakan komputer umum
    • Belanja pada situs yang ada pengamannya
    • Belajar untuk mendeteksi dan melaporkan segera aktivitas yang mencurigakan

Peliharalah hubungan komunikasi yang positif serta terbuka dengan anak-anak Anda. Anjurkan anak Anda untuk bicara dengan Anda jika ada keraguan mereka terhadap dampak Internet.

Wednesday, April 1, 2009

When the Internet Runs Out of IP Addresses

Source: http://hbswk.hbs.edu/item/5968.html

The Internet is running out of room.

Experts predict that in two or three years we will run out of Web addresses, so-called IP addresses, that can be assigned to new Internet-based sites and services.

Each site is assigned a unique number based on the IPv4 standard. IPv4 is the basis of addresses using combinations of four integers, about 4 billion possible combinations.

The problem: Internet growth is so dramatic that potential IPv4 addresses are running short on supply. A new standard called IPv6 will solve the problem with trillions of possible combinations, but technical and practical roadblocks are delaying its widespread implementation, probably for many years.

So what happens when the last IPv4 address is assigned? Harvard Business School professor Benjamin G. Edelman proposes a solution: Create a market for holders of previously assigned but unused addresses to sell or otherwise transfer them to new owners.

"It's unlikely that other networks would return their space for free—why would they?" says Edelman. "But if the price is right, they may be willing to transfer the space to someone who needs it more."

Our Q&A follows.

Sean Silverthorne: Why are we running out of IPv4 addresses?

Ben Edelman: In this respect, the Internet is a victim of its own success. The IP address system was never designed for the large, complicated, widely used Internet we enjoy today. We're fortunate in that we've gotten this far without having to upgrade the Internet's addressing fundamentals. But the time is coming when such upgrades will be unavoidable.

Usually when a resource becomes scarce, its price increases. But that hasn't happened here because IP addresses are provided at de minimis cost—a regulatory decision reflecting that IP addresses are ultimately just numbers, and that it's odd to have to pay for a number. With prices stuck very close to zero, and demand steady and growing, economic incentives invite exhaustion.

Q: What happens if nothing is done and we run out of addresses? Does the Internet stop growing?

A: If networks can't get new IP addresses, it will be much harder for them to grow. They have some options, like making more intensive use of the addresses they already have through address sharing (formally, network address translation, or NAT) or perhaps revisiting which addresses truly need IP addresses. (Does that laser printer actually need a globally unique address? Maybe it could make due with a local-only address.)

But these approaches have major challenges. For example, NAT impedes many kinds of applications, like videoconferencing and certain file transfers. Large-scale NAT will limit innovation, increase complexity, and ultimately make the Internet less useful than it could be.

Another important challenge is that if networks can't get new IP addresses, it will be harder to enter many technology businesses. Want to start a new ISP? Or run a service provider that hosts a large number of Web sites? Without ample IP addresses, it's difficult to get started in these businesses. Entry and potential entry are an important part of competition. We need to make sure new firms can easily begin operations so that existing providers can't hold customers hostage.

Q: Isn't a transition to Ipv6 supposed to solve this problem?

A: IPv6 offers important benefits. In particular, if all networks ran IPv6, there would be plenty of addresses for everyone, and we'd have no further address shortage.

But it's difficult to get from here to there. Right now, users all run IPv4 computers, Web sites all run IPv4 Web servers, and ISPs all offer IPv4 transit. Transition raises the question of who goes first. ISPs have hesitated to offer IPv6 service because there's not much demand. Users (and end-user networks like companies and universities) aren't demanding IPv6 because, at least for now, they still can get all the IPv4 addresses they need. Neither are Web sites demanding IPv6, because Web sites want to be reached by users, and users only run IPv4.

In fact, it's even worse than that. Moving to IPv6 too early brings extra costs. For one, there are the standard costs of learning how to set up IPv6, and allocating engineering time to handle all the details. But when a Web site supports both IPv6 and IPv4, some users will mistakenly try to reach the site by IPv6 because their computers and network cards are misconfigured. (For example, certain security enables IPv6 in order to install IPv6 security—which seems like a good idea but actually results in v6 being enabled when it shouldn't be.)

These problems don't affect that many users—measurements suggest a fraction of a percent. But that's enough. If you're Dell, would you want to turn on IPv6, with no current benefits and an immediate loss of a fraction of a percent of your users? It just doesn't make sense, especially when your competitors can stay with v4 alone and be reachable by 100 percent of users as always.

Q: You advocate a market-based solution, at least as a temporary measure. Can you describe the idea of transferring unused IP addresses?

A: The basic idea is that some networks have ended up with more IPv4 addresses than they need, while others have less than they need (or will need in the near future).

Why do some networks have more than they need? A variety of reasons: Some networks received extra-large allocations of 16-plus million addresses in the Internet's early days, when addresses looked abundant and when it was hard to subdivide address blocks into just the right sizes. Other networks have scaled back their IPv4 needs through a change of business focus, or even through bankruptcy.

Q: What are the benefits of a market-based plan? How much time would this buy us?

A: A market-based approach offers a real benefit to those who still need more IPv4 addresses after ordinary supplies run out. Rather than being told that no more IPv4 space is available, on any terms or at any price, these networks could offer payments to get v4 space from others. It's unlikely that other networks would return their space for free—why would they? But if the price is right, they may be willing to transfer the space to someone who needs it more.

So the core benefit is allocative efficiency, moving scarce resources to those who need them most.

But there are other benefits, too. By putting a positive price on IPv4 space, a market mechanism would remind current v4 users that their v4 space is valuable, and that they might want to try to vacate it, to the extent they can, perhaps by moving to IPv6. A market basically tells networks: "We will pay you to use v6 instead." That's a transition incentive quite different from anything we've seen to date. That's a transition incentive that just might work.

Q: Wouldn't speculators take advantage to drive up prices?

A: Some people are definitely worried about speculators. But I don't think they're going to be a big problem here.

For one thing, the proposed transfer rules are slated to require that an address recipient be a bona fide network that actually needs, and can use, IP addresses. The responsible regulators, like North America's American Registry for Internet Numbers (ARIN), have been making these determinations for more than a decade. They're going to keep performing these reviews even when the resources at issue are resources obtained via paid transfers, rather than resources obtained from as-yet-unused space.

Furthermore, the dynamics of this market will probably make it unappealing to speculators. In the long run, the Internet will move to IPv6. So anyone speculating on IPv4 knows the price is going to plummet, probably to zero, in due course.

How soon will that be? That's much harder to say. This is a one-of-a-kind transition, importantly different from the various other standards transitions we've faced before. For example, the digital TV transition benefited from a strong central authority, the U.S. government, that could order transition on a particular date certain. Not so here, for there's no one to order ISPs to run v6 rather than v4. So speculators would try to predict the future at their peril.

Q: If your plan was adopted, would consumers pay more for Internet access and to develop Web sites?

A: If this transition goes smoothly, consumers should never notice. To date, IP addresses have been a trivially small part of the cost of Internet access and Web site hosting. Even if IP address prices increased 100 times, consumers still probably wouldn't notice.

The bigger worries come if ISPs just cannot expand, or just cannot enter the market. If that were to come to pass, I wouldn't be surprised to see effects on service price and quality. That's why it's so important to make the transition smooth—to provide financial incentives to move to IPv6, and to provide a framework to let ISPs enter and expand even their IPv4 operations at reasonable cost and with appropriate predictability.

Q: Who will make this decision? Will governments play a role?

A: IP addresses are given out by five Regional Internet Registries (RIRs). In North America, our RIR is the American Registry for Internet Numbers (ARIN). RIRs are private nonprofits, not a government agency, and their powers are appropriately limited. But RIRs are in a position to allow paid transfers, if they conclude that such transfers are in the Internet's best interests.

Governments definitely play a role. For example, the U.S. Office of Management and Budget (OMB) mandated the installation of IPv6-capable equipment on agencies' network backbones by June 2008—certainly a step in the right direction, and one of the few major drivers of IPv6 deployment. But so far OMB hasn't yet tried to force agencies to actually use IPv6, knowing that it would be too costly, too inconvenient, and on balance too hard to justify, at least for now.

More generally, the Internet's infrastructure is largely private. ISPs make their own decisions about what systems to install, and what services to provide. Certainly governments can offer incentives, but when Japan experimented with IPv6 deployment incentives at the start of the decade, the payments had limited effectiveness.

So at present, I think the Internet's best hopes come not from governments or outside deployment incentives, but from internal incentives grounded in putting a positive price on increasingly valuable IPv4 resources.

Friday, June 27, 2008

Caching Tutorial

What's a Web Cache? Why do people use them?



A Web cache sits between one or more Web servers (also known as
origin servers) and a client or many clients, and watches requests
come by, saving copies of the responses : like HTML pages, images and files
(collectively known as representations) : for itself. Then, if there
is another request for the same URL, it can use the response that it has,
instead of asking the origin server for it again.



There are two main reasons that Web caches are used:




  • To reduce latency : Because the request is satisfied
    from the cache (which is closer to the client) instead of the origin server,
    it takes less time for it to get the representation and display it. This
    makes the Web seem more responsive.

  • To reduce network traffic : Because representations are
    reused, it reduces the amount of bandwidth used by a client. This saves
    money if the client is paying for traffic, and keeps their bandwidth
    requirements lower and more manageable.



Kinds of Web Caches



Browser Caches



If you examine the preferences dialog of any modern Web browser (like
Internet Explorer, Safari or Mozilla), you'll probably notice a cache
setting. This lets you set aside a section of your computer's hard disk to
store representations that you've seen, just for you. The browser cache works
according to fairly simple rules. It will check to make sure that the
representations are fresh, usually once a session (that is, the once in the
current invocation of the browser).



This cache is especially useful when users hit the back button or click a
link to see a page they've just looked at. Also, if you use the same
navigation images throughout your site, they'll be served from browsers'
caches almost instantaneously.




Proxy Caches



Web proxy caches work on the same principle, but a much larger scale.
Proxies serve hundreds or thousands of users in the same way; large
corporations and ISPs often set them up on their firewalls, or as standalone
devices (also known as intermediaries).



Because proxy caches aren't part of the client or the origin server, but
instead are out on the network, requests have to be routed to them somehow.
One way to do this is to use your browser's proxy setting to manually tell it
what proxy to use; another is using interception. Interception
proxies
have Web requests redirected to them by the underlying
network itself, so that clients don't need to be configured for them, or even
know about them.



Proxy caches are a type of shared cache; rather than just having
one person using them, they usually have a large number of users, and because
of this they are very good at reducing latency and network traffic. That's
because popular representations are reused a number of times.




Gateway Caches



Also known as reverse proxy caches or surrogate caches, gateway caches
are also intermediaries, but instead of being deployed by network
administrators to save bandwidth, they're typically deployed by Webmasters
themselves, to make their sites more scalable, reliable and better
performing.



Requests can be routed to gateway caches by a number of methods, but
typically some form of load balancer is used to make one or more of them look
like the origin server to clients.



Application delivery networks (ADNs) distribute gateway caches
throughout the Internet (or a part of it) and sell caching to interested Web
sites. Speedera and Akamai are examples of
ADNs.



This tutorial focuses mostly on browser and proxy caches, although some of
the information is suitable for those interested in gateway caches as
well.




Aren't Web Caches bad for me? Why should I help
them?



Web caching is one of the most misunderstood technologies on the Internet.
Webmasters in particular fear losing control of their site, because a proxy
cache can hide their users from them, making it difficult to see who's using
the site.



Unfortunately for them, even if Web caches didn't exist, there are too many
variables on the Internet to assure that they'll be able to get an accurate
picture of how users see their site. If this is a big concern for you, this
tutorial will teach you how to get the statistics you need without making your
site cache-unfriendly.



Another concern is that caches can serve content that is out of date, or
stale. However, this tutorial can show you how to configure your
server to control how your content is cached.



ADNs
are an interesting development, because unlike many
proxy caches, their gateway caches are aligned with the interests of the
Web site being cached, so that these problems aren't seen. However, even
when you use a ADN, you still have to consider that there will be proxy
and browser caches downstream.



On the other hand, if you plan your site well, caches can help your Web
site load faster, and save load on your server and Internet link. The
difference can be dramatic; a site that is difficult to cache may take
several seconds to load, while one that takes advantage of caching can seem
instantaneous in comparison. Users will appreciate a fast-loading site, and
will visit more often.



Think of it this way; many large Internet companies are spending millions
of dollars setting up farms of servers around the world to replicate their
content, in order to make it as fast to access as possible for their users.
Caches do the same for you, and they're even closer to the end user. Best of
all, you don't have to pay for them.



The fact is that proxy and browser caches will be used whether you like it
or not. If you don't configure your site to be cached correctly, it will be
cached using whatever defaults the cache's administrator decides upon.




How Web Caches Work



All caches have a set of rules that they use to determine when to serve a
representation from the cache, if it's available. Some of these rules are set
in the protocols (HTTP 1.0 and 1.1), and some are set by the administrator of
the cache (either the user of the browser cache, or the proxy
administrator).



Generally speaking, these are the most common rules that are followed
(don't worry if you don't understand the details, it will be explained
below):





  1. If the response's headers tell the cache not to keep it,
    it won't.

  2. If the request is authenticated or secure, it won't be
    cached.

  3. If no validator (an ETag or Last-Modified header) is
    present on a response, and it doesn't have any explicit freshness information,
    it will be considered uncacheable.

  4. A cached representation is considered fresh (that is, able to
    be sent to a client without checking with the origin server) if:

    • It has an expiry time or other age-controlling header set, and is
      still within the fresh period.

    • If a browser cache has already seen the representation, and has been
      set to check once a session.

    • If a proxy cache has seen the representation recently, and it was
      modified relatively long ago.


    Fresh representations are served directly from the cache, without checking
    with the origin server.

  5. If an representation is stale, the origin server will be asked to
    validate it, or tell the cache whether the copy that it has is
    still good.




Together, freshness and validation are the most important
ways that a cache works with content. A fresh representation will be available
instantly from the cache, while a validated representation will avoid sending
the entire representation over again if it hasn't changed.



How (and how not) to Control
Caches



There are several tools that Web designers and Webmasters can use to
fine-tune how caches will treat their sites. It may require getting your hands
a little dirty with your server's configuration, but the results are worth it.
For details on how to use these tools with your server, see the Implementation sections below.



HTML Meta Tags and HTTP Headers



HTML authors can put tags in a document's <HEAD> section that
describe its attributes. These meta tags are often used in the
belief that they can mark a document as uncacheable, or expire it at a
certain time.



Meta tags are easy to use, but aren't very effective. That's because
they're only honored by a few browser caches (which actually read the HTML),
not proxy caches (which almost never read the HTML in the document). While it
may be tempting to put a Pragma: no-cache meta tag into a Web page, it won't
necessarily cause it to be kept fresh.



If your site is hosted at an ISP or hosting farm and they
don't give you the ability to set arbitrary HTTP headers (like Expires and
Cache-Control), complain loudly; these are tools necessary for doing your
job.



On the other hand, true HTTP headers give you a lot of control
over how both browser caches and proxies handle your representations. They
can't be seen in the HTML, and are usually automatically generated by the Web
server. However, you can control them to some degree, depending on the server
you use. In the following sections, you'll see what HTTP headers are
interesting, and how to apply them to your site.



HTTP headers are sent by the server before the HTML, and only seen by the
browser and any intermediate caches. Typical HTTP 1.1 response headers might
look like this:



HTTP/1.1  OK
Date: Fri, 30 Oct 1998 13:19:41 GMT
Server: Apache/1.3.3 (Unix)
Cache-Control: max-age=3600, must-revalidate
Expires: Fri, 30 Oct 1998 14:19:41 GMT
Last-Modified: Mon, 29 Jun 1998 02:28:12 GMT
ETag: "3e86-410-3596fbbc"
Content-Length: 1040
Content-Type: text/html


The HTML would follow these headers, separated by a blank
line. See the Implementation sections for information about how to set HTTP
headers.





Pragma HTTP Headers (and why they don't
work)



Many people believe that assigning a Pragma: no-cache HTTP header to a
representation will make it uncacheable. This is not necessarily true; the
HTTP specification does not set any guidelines for Pragma response headers;
instead, Pragma request headers (the headers that a browser sends to a server)
are discussed. Although a few caches may honor this header, the majority
won't, and it won't have any effect. Use the headers below instead.



Controlling Freshness with the Expires
HTTP Header



The Expires HTTP header is a basic means of controlling caches; it tells
all caches how long the associated representation is fresh for. After that
time, caches will always check back with the origin server to see if a
document is changed. Expires headers are supported by practically every
cache.



Most Web servers allow you to set Expires response headers in a number of
ways. Commonly, they will allow setting an absolute time to expire, a time
based on the last time that the client saw the representation (last access
time
), or a time based on the last time the document changed on your
server (last modification time).



Expires headers are especially good for making static images (like
navigation bars and buttons) cacheable. Because they don't change much, you
can set extremely long expiry time on them, making your site appear much more
responsive to your users. They're also useful for controlling caching of a
page that is regularly changed. For instance, if you update a news page once a
day at 6am, you can set the representation to expire at that time, so caches
will know when to get a fresh copy, without users having to hit ‘reload'.



The only value valid in an Expires header is a HTTP date;
anything else will most likely be interpreted as ‘in the past', so that the
representation is uncacheable. Also, remember that the time in a HTTP date is
Greenwich Mean Time (GMT), not local time.



For example:


Expires: Fri, 30 Oct 1998 14:19:41 GMT


It's important to make sure that your Web
server's clock is accurate if you use the Expires header.
One way to do this is using the Network Time
Protocol
(NTP); talk to your local system administrator to find out
more.



Although the Expires header is useful, it has some limitations. First,
because there's a date involved, the clocks on the Web server and the cache
must be synchronised; if they have a different idea of the time, the intended
results won't be achieved, and caches might wrongly consider stale content as
fresh.




Another problem with Expires is that it's easy to forget that you've set
some content to expire at a particular time. If you don't update an Expires
time before it passes, each and every request will go back to your Web server,
increasing load and latency.



Cache-Control HTTP
Headers



HTTP 1.1 introduced a new class of headers, Cache-Control response
headers, to give Web publishers more control over their content, and
to address the limitations of Expires.



Useful Cache-Control response headers include:




  • max-age=[seconds] : specifies the maximum amount of
    time that an representation will be considered fresh. Similar to Expires,
    this directive is relative to the time of the request, rather than absolute.
    [seconds] is the number of seconds from the time of the request you wish the
    representation to be fresh for.

  • s-maxage=[seconds] : similar to max-age, except that it
    only applies to shared (e.g., proxy) caches.

  • public : marks authenticated responses as cacheable;
    normally, if HTTP authentication is required, responses are automatically
    uncacheable.

  • no-cache : forces caches to submit the request to the
    origin server for validation before releasing a cached copy, every time.
    This is useful to assure that authentication is respected (in combination
    with public), or to maintain rigid freshness, without sacrificing all of the
    benefits of caching.

  • no-store : instructs caches not to keep a copy of the
    representation under any conditions.

  • must-revalidate : tells caches that they must obey any
    freshness information you give them about a representation. HTTP allows
    caches to serve stale representations under special conditions; by
    specifying this header, you're telling the cache that you want it to
    strictly follow your rules.

  • proxy-revalidate : similar to must-revalidate, except
    that it only applies to proxy caches.



For example:


Cache-Control: max-age=3600, must-revalidate


If you plan to use the Cache-Control headers, you should have a look at
the excellent documentation in HTTP 1.1; see References and Further Information.



Validators and Validation



In How Web Caches Work, we said that validation is used
by servers and caches to communicate when an representation has changed. By
using it, caches avoid having to download the entire representation when they
already have a copy locally, but they're not sure if it's still fresh.



Validators are very important; if one isn't present, and there isn't any
freshness information (Expires or Cache-Control) available, caches will
not store a representation at all.



The most common validator is the time that the document last changed, as
communicated in Last-Modified header. When a cache has an
representation stored that includes a Last-Modified header, it can use it to
ask the server if the representation has changed since the last time it was
seen, with an If-Modified-Since request.



HTTP 1.1 introduced a new kind of validator called the ETag. ETags
are unique identifiers that are generated by the server and changed every time
the representation does. Because the server controls how the ETag is
generated, caches can be surer that if the ETag matches when they make a
If-None-Match request, the representation really is the same.



Almost all caches use Last-Modified times in determining if an
representation is fresh; ETag validation is also becoming prevalent.



Most modern Web servers will generate both ETag and Last-Modified
headers to use as validators for static content (i.e., files) automatically; you won't have to
do anything. However, they don't know enough about dynamic content (like CGI,
ASP or database sites) to generate them; see Writing
Cache-Aware Scripts
.



Tips for Building a Cache-Aware Site



Besides using freshness information and validation, there are a number of
other things you can do to make your site more cache-friendly.



  • Use URLs consistently : this is the golden
    rule of caching. If you serve the same content on different pages, to
    different users, or from different sites, it should use the same URL.
    This is the easiest and most effective may to make your site
    cache-friendly. For example, if you use /index.html in your HTML as a
    reference once, always use it that way.

  • Use a common library of images and other elements and
    refer back to them from different places.

  • Make caches store images and pages that don't change
    often
    by using a Cache-Control: max-age header with a large
    value.

  • Make caches recognize regularly updated pages by
    specifying an appropriate max-age or expiration time.

  • If a resource (especially a downloadable file) changes, change
    its name.
    That way, you can make it expire far in the future,
    and still guarantee that the correct version is served; the page that
    links to it is the only one that will need a short expiry time.

  • Don't change files unnecessarily. If you do,
    everything will have a falsely young Last-Modified date. For instance,
    when updating your site, don't copy over the entire site; just move the
    files that you've changed.

  • Use cookies only where necessary : cookies are
    difficult to cache, and aren't needed in most situations. If you must use
    a cookie, limit its use to dynamic pages.

  • Minimize use of SSL : because encrypted pages are not
    stored by shared caches, use them only when you have to, and use images
    on SSL pages sparingly.

  • use the Cacheability Engine
    : it can help you apply many of the concepts in this tutorial.



Writing Cache-Aware Scripts



By default, most scripts won't return a validator (a Last-Modified
or ETag response header) or freshness information (Expires or Cache-Control).
While some scripts really are dynamic (meaning that they return a different
response for every request), many (like search engines and database-driven
sites) can benefit from being cache-friendly.



Generally speaking, if a script produces output that is reproducable with
the same request at a later time (whether it be minutes or days later), it
should be cacheable. If the content of the script changes only depending on
what's in the URL, it is cacheble; if the output depends on a cookie,
authentication information or other external criteria, it probably isn't.



  • The best way to make a script cache-friendly (as well as perform
    better) is to dump its content to a plain file whenever it changes. The
    Web server can then treat it like any other Web page, generating and
    using validators, which makes your life easier. Remember to only write
    files that have changed, so the Last-Modified times are preserved.

  • Another way to make a script cacheable in a limited fashion is to set
    an age-related header for as far in the future as practical. Although
    this can be done with Expires, it's probably easiest to do so with
    Cache-Control: max-age, which will make the request fresh for an amount
    of time after the request.

  • If you can't do that, you'll need to make the script generate a
    validator, and then respond to If-Modified-Since and/or If-None-Match
    requests. This can be done by parsing the HTTP headers, and then
    responding with 304 Not Modified when appropriate. Unfortunately, this is
    not a trival task.



Some other tips;



  • Don't use POST unless it's appropriate. Responses to
    the POST method aren't kept by most caches; if you send information in the
    path or query (via GET), caches can store that information for the
    future.

  • Don't embed user-specific information in the URL unless
    the content generated is completely unique to that user.

  • Don't count on all requests from a user coming from the same
    host
    , because caches often work together.

  • Generate Content-Length response headers. It's easy to
    do, and it will allow the response of your script to be used in a
    persistent connection. This allows clients to request
    multiple representations on one TCP/IP connection, instead of setting up a
    connection for every request. It makes your site seem much faster.



See the Implementation Notes for more specific
information.



Frequently Asked Questions



What are the most important things to make cacheable?



A good strategy is to identify the most popular, largest representations
(especially images) and work with them first.



How can I make my pages as fast as possible with caches?



The most cacheable representation is one with a long freshness time set.
Validation does help reduce the time that it takes to see a representation,
but the cache still has to contact the origin server to see if it's fresh. If
the cache already knows it's fresh, it will be served directly.



I understand that caching is good, but I need to keep statistics on how
many people visit my page!



If you must know every time a page is accessed, select ONE small item on
a page (or the page itself), and make it uncacheable, by giving it a suitable
headers. For example, you could refer to a 1x1 transparent uncacheable image
from each page. The Referer header will contain information about what page
called it.



Be aware that even this will not give truly accurate statistics about your
users, and is unfriendly to the Internet and your users; it generates
unnecessary traffic, and forces people to wait for that uncached item to be
downloaded. For more information about this, see On Interpreting Access
Statistics in the references.



How can I see a representation's HTTP headers?



Many Web browsers let you see the Expires and Last-Modified headers are in
a page info or similar interface. If available, this will give you a menu of
the page and any representations (like images) associated with it, along with
their details.



To see the full headers of a representation, you can manually connect to
the Web server using a Telnet client.



To do so, you may need to type the port (be default, 80) into a separate
field, or you may need to connect to www.example.com:80 or www.example.com 80
(note the space). Consult your Telnet client's documentation.



Once you've opened a connection to the site, type a request for the
representation. For instance, if you want to see the headers for
http://www.example.com/foo.html, connect to www.example.com, port 80, and
type:


GET /foo.html HTTP/1.1 [return]
Host: www.example.com [return][return]


Press the Return key every time you see [return]; make sure to press it
twice at the end. This will print the headers, and then the full
representation. To see the headers only, substitute HEAD for GET.




My pages are password-protected; how do proxy caches deal with them?



By default, pages protected with HTTP authentication are considered private;
they will not be kept by shared caches. However, you can make authenticated
pages public with a Cache-Control: public header; HTTP 1.1-compliant caches will then
allow them to be cached.



If you'd like such pages to be cacheable, but still authenticated for every
user, combine the Cache-Control: public and no-cache headers. This tells the
cache that it must submit the new client's authentication information to the
origin server before releasing the representation from the cache. This would look like:



Cache-Control: public, no-cache


Whether or not this is done, it's best to minimize use of authentication;
for example, if your images are not sensitive, put them in a separate
directory and configure your server not to force authentication for it. That
way, those images will be naturally cacheable.




Should I worry about security if people access my site through a
cache?



SSL pages are not cached (or decrypted) by proxy caches, so you don't have
to worry about that. However, because caches store non-SSL requests and URLs
fetched through them, you should be conscious about unsecured sites; an
unscrupulous administrator could conceivably gather information about their
users, especially in the URL.



In fact, any administrator on the network between your server and your
clients could gather this type of information. One particular problem is when
CGI scripts put usernames and passwords in the URL itself; this makes it
trivial for others to find and user their login.



If you're aware of the issues surrounding Web security in general, you
shouldn't have any surprises from proxy caches.



I'm looking for an integrated Web publishing solution. Which ones are
cache-aware?



It varies. Generally speaking, the more complex a solution is, the more
difficult it is to cache. The worst are ones which dynamically generate all
content and don't provide validators; they may not be cacheable at all. Speak
with your vendor's technical staff for more information, and see the
Implementation notes below.



My images expire a month from now, but I need to change them in the
caches now!



The Expires header can't be circumvented; unless the cache (either browser
or proxy) runs out of room and has to delete the representations, the cached
copy will be used until then.



The most effective solution is to change any links to them; that way,
completely new representations will be loaded fresh from the origin server.
Remember that the page that refers to an representation will be cached as
well. Because of this, it's best to make static images and similar
representations very cacheable, while keeping the HTML pages that refer to
them on a tight leash.



If you want to reload an representation from a specific cache, you can
either force a reload (in Firefox, holding down shift while pressing ‘reload'
will do this by issuing a Pragma: no-cache request header) while using the
cache. Or, you can have the cache administrator delete the representation
through their interface.



I run a Web Hosting service. How can I let my users publish
cache-friendly pages?



If you're using Apache, consider allowing them to use .htaccess files and
providing appropriate documentation.



Otherwise, you can establish predetermined areas for various caching
attributes in each virtual server. For instance, you could specify a
directory /cache-1m that will be cached for one month after access, and a
/no-cache area that will be served with headers instructing caches not to
store representations from it.



Whatever you are able to do, it is best to work with your largest
customers first on caching. Most of the savings (in bandwidth and in load on
your servers) will be realized from high-volume sites.



I've marked my pages as cacheable, but my browser keeps requesting them
on every request. How do I force the cache to keep representations of them?



Caches aren't required to keep a representation and reuse it; they're only
required to not keep or use them under some conditions. All
caches make decisions about which representations to keep based upon their
size, type (e.g., image vs. html), or by how much space they have left to keep
local copies. Yours may not be considered worth keeping around, compared to
more popular or larger representations.



Some caches do allow their administrators to prioritize what kinds of
representations are kept, and some allow representations to be pinned in
cache, so that they're always available.




Implementation Notes : Web
Servers



Generally speaking, it's best to use the latest version of whatever Web
server you've chosen to deploy. Not only will they likely contain more
cache-friendly features, new versions also usually have important security
and performance improvements.



Apache HTTP Server



Apache uses
optional modules to include headers, including both Expires and
Cache-Control. Both modules are available in the 1.2 or greater
distribution.



The modules need to be built into Apache; although they are included in
the distribution, they are not turned on by default. To find out if the
modules are enabled in your server, find the httpd binary and run httpd
-l
; this should print a list of the available modules. The modules we're
looking for are mod_expires and mod_headers.



  • If they aren't available, and you have administrative access, you can
    recompile Apache to include them. This can be done either by uncommenting
    the appropriate lines in the Configuration file, or using the
    -enable-module=expires and -enable-module=headers
    arguments to configure (1.3 or greater). Consult the INSTALL file found
    with the Apache distribution.



Once you have an Apache with the appropriate modules, you can use
mod_expires to specify when representations should expire, either in .htaccess
files or in the server's access.conf file. You can specify expiry from either
access or modification time, and apply it to a file type or as a default. See
the module
documentation
for more information, and speak with your local Apache guru
if you have trouble.



To apply Cache-Control headers, you'll need to use the mod_headers module,
which allows you to specify arbitrary HTTP headers for a resource. See the
mod_headers documentation
.



Here's an example .htaccess file that demonstrates the use of some
headers.



  • .htaccess files allow web publishers to use commands normally only
    found in configuration files. They affect the content of the directory
    they're in and their subdirectories. Talk to your server administrator to
    find out if they're enabled.


### activate mod_expires
ExpiresActive On
### Expire .gif's 1 month from when they're accessed
ExpiresByType image/gif A2590
### Expire everything else 1 day from when it's last modified
### (this uses the Alternative syntax)
ExpiresDefault "modification plus 1 day"
### Apply a Cache-Control header to index.html
<Files index.html>
Header append Cache-Control "public, must-revalidate"
</Files>


  • Note that mod_expires automatically calculates and inserts a
    Cache-Control:max-age header as appropriate.



Apache 2.0's configuration is very similar to that of 1.3; see the 2.0 mod_expires and
mod_headers
documentation for more information.




Microsoft IIS



Microsoft's
Internet Information Server makes it very easy to set headers in a somewhat
flexible way. Note that this is only possible in version 4 of the server,
which will run only on NT Server.



To specify headers for an area of a site, select it in the
Administration Tools interface, and bring up its properties. After
selecting the HTTP Headers tab, you should see two interesting
areas; Enable Content Expiration and Custom HTTP headers.
The first should be self-explanatory, and the second can be used to apply
Cache-Control headers.



See the ASP section below for information about setting headers in Active
Server Pages. It is also possible to set headers from ISAPI modules; refer to
MSDN for details.




Netscape/iPlanet Enterprise Server



As of version 3.6, Enterprise Server does not provide any obvious way to
set Expires headers. However, it has supported HTTP 1.1 features since version
3.0. This means that HTTP 1.1 caches (proxy and browser) will be able to take
advantage of Cache-Control settings you make.



To use Cache-Control headers, choose Content Management | Cache Control
Directives
in the administration server. Then, using the Resource Picker,
choose the directory where you want to set the headers. After setting the
headers, click ‘OK'. For more information, see the NES manual.




Implementation Notes : Server-Side
Scripting



One thing to keep in mind is that it may be easier to set
HTTP headers with your Web server rather than in the scripting language. Try
both.



Because the emphasis in server-side scripting is on dynamic content, it
doesn't make for very cacheable pages, even when the content could be cached.
If your content changes often, but not on every page hit, consider setting a
Cache-Control: max-age header; most users access pages again in a relatively
short period of time. For instance, when users hit the ‘back' button, if there
isn't any validator or freshness information available, they'll have to wait
until the page is re-downloaded from the server to see it.




CGI



CGI scripts are one of the most popular ways to generate content. You can
easily append HTTP response headers by adding them before you send the body;
Most CGI implementations already require you to do this for the
Content-Type header. For instance, in Perl;



#!/usr/bin/perl
print "Content-type: text/html\n";
print "Expires: Thu, 29 Oct 1998 17:04:19 GMT\n";
print "\n";
### the content body follows...


Since it's all text, you can easily generate Expires and other
date-related headers with in-built functions. It's even easier if you use
Cache-Control: max-age;



print "Cache-Control: max-age=600\n";


This will make the script cacheable for 10 minutes after the request, so
that if the user hits the ‘back' button, they won't be resubmitting the
request.



The CGI specification also makes request headers that the client sends
available in the environment of the script; each header has ‘HTTP_' appended
to its name. So, if a client makes an If-Modified-Since request, it may show
up like this:



HTTP_IF_MODIFIED_SINCE = Fri, 30 Oct 1998 14:19:41 GMT 


See also the cgi_buffer
library, which automatically handles ETag generation and validation,
Content-Length generation and gzip content-oding for Perl and Python CGI
scripts with a one-line include. The Python version can also be used to wrap
arbitrary CGI scripts with.



Server Side Includes



SSI (often used with the extension .shtml) is one of the first ways that
Web publishers were able to get dynamic content into pages. By using special
tags in the pages, a limited form of in-HTML scripting was available.



Most implementations of SSI do not set validators, and as such are not
cacheable. However, Apache's implementation does allow users to specify which
SSI files can be cached, by setting the group execute permissions on the
appropriate files, combined with the XbitHack full directive. For more
information, see the mod_include
documentation
.



PHP



PHP is a
server-side scripting language that, when built into the server, can be used
to embed scripts inside a page's HTML, much like SSI, but with a far larger
number of options. PHP can be used as a CGI script on any Web server (Unix or
Windows), or as an Apache module.



By default, representations processed by PHP are not assigned validators,
and are therefore uncacheable. However, developers can set HTTP headers by
using the Header() function.



For example, this will create a Cache-Control header, as well as an
Expires header three days in the future:


<?php
Header("Cache-Control: must-revalidate");

$offset = 60 * 60 * 24 * 3;
$ExpStr = "Expires: " . gmdate("D, d M Y H:i:s", time() + $offset) . " GMT";
Header($ExpStr);
?>


Remember that the Header() function MUST come before any other output.



As you can see, you'll have to create the HTTP date for an Expires header
by hand; PHP doesn't provide a function to do it for you (although recent
versions have made it easier; see the PHP's date documentation). Of course, it's
easy to set a Cache-Control: max-age header, which is just as good for most
situations.



For more information, see the manual entry for
header
.



See also the cgi_buffer library, which
automatically handles ETag generation and validation, Content-Length
generation and gzip content-coding for PHP scripts with a one-line
include.



Cold Fusion



Macromedia is a commercial server-side
scripting engine, with support for several Web servers on Windows, Linux and
several flavors of Unix.



Cold Fusion makes setting arbitrary HTTP headers relatively easy, with the
CFHEADER
tag. Unfortunately, their example for setting an Expires header, as below, is a bit misleading.





It doesn't work like you might think, because the time (in this case, when the request is made)
doesn't get converted to a HTTP-valid date; instead, it just gets printed as
a representation of Cold Fusion's Date/Time object. Most clients will either
ignore such a value, or convert it to a default, like January 1, 1970.



However, Cold Fusion does provide a date formatting function that will do the job;
GetHttpTimeSTring. In combination with
DateAdd
, it's easy to set Expires dates;
here, we set a header to declare that representations of the page expire in one month;





You can also use the CFHEADER tag to set Cache-Control: max-age and other headers.



Remember that Web server headers are passed through in some deployments of Cold Fusion
(such as CGI); check yours to determine whether you can use
this to your advantage, by setting headers on the server instead of in Cold
Fusion.



ASP and ASP.NET



When setting HTTP headers from ASPs, make sure you either
place the Response method calls before any HTML generation, or use
Response.Buffer to buffer the output. Also, note that some versions of IIS set
a Cache-Control: private header on ASPs by default, and must be declared public
to be cacheable by shared caches.



Active Server Pages, built into IIS and also available for other Web
servers, also allows you to set HTTP headers. For instance, to set an expiry
time, you can use the properties of the Response object;



<% Response.Expires=1440 %>


specifying the number of minutes from the request to expire the
representation. Likewise, absolute expiry time can be set like this (make sure
you format HTTP date correctly):



<% Response.ExpiresAbsolute=#May 31,1996 13:30:15 GMT# %>


Cache-Control headers can be added like this:






In ASP.NET, Response.Expires is deprecated; the proper way to set cache-related
headers is with Response.Cache;



Response.Cache.SetExpires ( DateTime.Now.AddMinutes ( 60 ) ) ;
Response.Cache.SetCacheability ( HttpCacheability.Public ) ;


See the MSDN documentation for more
information.




References and Further Information



HTTP 1.1 Specification



The HTTP 1.1 spec has many extensions for making pages cacheable,
and is the authoritative guide to implementing the protocol. See sections 13,
14.9, 14.21, and 14.25.



Web-Caching.com



An excellent introduction to caching concepts, with links to other online
resources.



On Interpreting
Access Statistics



Jeff Goldberg's informative rant on why you shouldn't rely on access
statistics and hit counters.



Cacheability Engine



Examines Web pages to determine how they will interact with Web caches,
the Engine is a good debugging tool, and a companion to this tutorial.



cgi_buffer Library



One-line include in Perl CGI, Python CGI and PHP scripts automatically
handles ETag generation and validation, Content-Length generation and gzip
Content-Encoding : correctly. The Python version can also be used as a
wrapper around arbitrary CGI scripts.

Sunday, May 11, 2008

FTP (File Transfer Protocols)

File Transfer Protocol (FTP) is a network protocol used to transfer data from one computer to another through a network, such as over the Internet. FTP is a file transfer protocol for exchanging files over any TCP/IP based network to manipulate files on another computer on that network regardless of which operating systems are involved (if the computers permit FTP access). There are many existing FTP client and server programs. FTP servers can be set up anywhere between game servers, voice servers, internet hosts, and other physical servers.


Connections Methods

FTP runs exclusively over TCP. FTP servers by default listen on port 21 for incoming connections from FTP clients. A connection to this port from the FTP Client forms the control stream on which commands are passed to the FTP server from the FTP client and on occasion from the FTP server to the FTP client. FTP uses out-of-band control, which means it uses a separate connection for control and data. Thus, for the actual file transfer to take place, a different connection is required which is called the data stream. Depending on the transfer mode, the process of setting up the data stream is different.
In active mode, the FTP client opens a dynamic port (49152–65535), sends the FTP server the dynamic port number on which it is listening over the control stream and waits for a connection from the FTP server. When the FTP server initiates the data connection to the FTP client it binds the source port to port 20 on the FTP server. In order to use active mode, the client sends a PORT command, with the IP and port as argument. The format for the IP and port is "h1,h2,h3,h4,p1,p2". Each field is a decimal representation of 8 bits of the host IP, followed by the chosen data port. For example, a client with an IP of 192.168.0.1, listening on port 49154 for the data connection will send the command "PORT 192,168,0,1,192,2". The port fields should be interpreted as p1×256 + p2 = port, or, in this example, 192×256 + 2 = 49154. In passive mode, the FTP server opens a dynamic port (49152–65535), sends the FTP client the server's IP address to connect to and the port on which it is listening (a 16 bit value broken into a high and low byte, like explained before) over the control stream and waits for a connection from the FTP client. In this case the FTP client binds the source port of the connection to a dynamic port between 49152 and 65535. To use passive mode, the client sends the PASV command to which the server would reply with something similar to "227 Entering Passive Mode (127,0,0,1,192,52)". The syntax of the IP address and port are the same as for the argument to the PORT command. In extended passive mode, the FTP server operates exactly the same as passive mode, however it only transmits the port number (not broken into high and low bytes) and the client is to assume that it connects to the same IP address that was originally connected to. Extended passive mode was added by RFC 2428 in September 1998. While data is being transferred via the data stream, the control stream sits idle. This can cause problems with large data transfers through firewalls which time out sessions after lengthy periods of idleness. While the file may well be successfully transferred, the control session can be disconnected by the firewall, causing an error to be generated. The FTP protocol supports resuming of interrupted downloads using the REST command. The client passes the number of bytes it has already received as argument to the REST command and restarts the transfer. In some commandline clients for example, there is an often-ignored but valuable command, "reget" (meaning "get again") that will cause an interrupted "get" command to be continued, hopefully to completion, after a communications interruption. Resuming uploads is not as easy. Although the FTP protocol supports the APPE command to append data to a file on the server, the client does not know the exact position at which a transfer got interrupted. It has to obtain the size of the file some other way, for example over a directory listing or using the SIZE command. In ASCII mode (see below), resuming transfers can be troublesome if client and server use different end of line characters. The objectives of FTP, as outlined by its RFC, are:
  1. To promote sharing of files (computer programs and/or data).
  2. To encourage indirect or implicit use of remote computers.
  3. To shield a user from variations in file storage systems among different hosts.
  4. To transfer data reliably, and efficiently.

Criticisms of FTP
  • Passwords and file contents are sent in clear text, which can be intercepted by eavesdroppers. There are protocol enhancements that circumvent this, for instance by using SSL, TLS or Kerberos.
  • Multiple TCP/IP connections are used, one for the control connection, and one for each download, upload, or directory listing. Firewalls may need additional logic and or configuration changes to account for these connections.
  • It is hard to filter active mode FTP traffic on the client side by using a firewall, since the client must open an arbitrary port in order to receive the connection. This problem is largely resolved by using passive mode FTP.
  • It is possible to abuse the protocol's built-in proxy features to tell a server to send data to an arbitrary port of a third computer; see FXP.
  • FTP is a high latency protocol due to the number of commands needed to initiate a transfer.
  • No integrity check on the receiver side. If a transfer is interrupted, the receiver has no way to know if the received file is complete or not. Some servers support extensions to calculate for example a file's MD5 sum (e.g. using the SITE MD5 command) or CRC checksum, however even then the client has to make explicit use of them. In the absence of such extensions, integrity checks have to be managed externally.
  • No date/timestamp attribute transfer. Uploaded files are given a new current timestamp, unlike other file transfer protocols such as SFTP, which allow attributes to be included. There is no way in the standard FTP protocol to set the time-last-modified (or time-created) datestamp that most modern filesystems preserve. There is a draft of a proposed extension that adds new commands for this, but as of yet, most of the popular FTP servers do not support it.
Security Problems
The original FTP specification is an inherently insecure method of transferring files because there is no method specified for transferring data in an encrypted fashion. This means that under most network configurations, user names, passwords, FTP commands and transferred files can be "sniffed" or viewed by anyone on the same network using a packet sniffer. This is a problem common to many Internet protocol specifications written prior to the creation of SSL such as HTTP, SMTP and Telnet. The common solution to this problem is to use either SFTP (SSH File Transfer Protocol), or FTPS (FTP over SSL), which adds SSL or TLS encryption to FTP as specified in RFC 4217

FTP Return Code
FTP server return codes indicate their status by the digits within them. A brief explanation of various digits' meanings are given below:
  • 1xx: Positive Preliminary reply. The action requested is being initiated but there will be another reply before it begins.
  • 2xx: Positive Completion reply. The action requested has been completed. The client may now issue a new command.
  • 3xx: Positive Intermediate reply. The command was successful, but a further command is required before the server can act upon the request.
  • 4xx: Transient Negative Completion reply. The command was not successful, but the client is free to try the command again as the failure is only temporary.
  • 5xx: Permanent Negative Completion reply. The command was not successful and the client should not attempt to repeat it again.
  • x0x: The failure was due to a syntax error.
  • x1x: This response is a reply to a request for information.
  • x2x: This response is a reply relating to connection information.
  • x3x: This response is a reply relating to accounting and authorization.
  • x4x: Unspecified as yet
  • x5x: These responses indicate the status of the Server file system vis-a-vis the requested transfer or other file system action.
Anonymous FTP
Many sites that run FTP servers enable anonymous ftp. Under this arrangement, users do not need an account on the server. The user name for anonymous access is typically 'anonymous', but historically 'ftp' was also used in the past; this account does not need a password.
Although users are commonly asked to send their email addresses as their passwords for "authentication," there is usually only trivial or no verification of what is actually entered. As modern FTP clients hide the login process from the user, and usually don't know the user's email address, the software supplies dummy passwords. For example: * Mozilla Firefox (2.0) — mozilla@example.com * KDE Konqueror (3.5) — anonymous@ * wget (1.10.2) — -wget@ * lftp (3.4.4) — lftp@ Internet Gopher has been suggested as an alternative to anonymous FTP, as well as Trivial File Transfer Protocol and File Service Protocol.

Data format
While transferring data over the network, several data representations can be used. The two most common transfer modes are:

  1. ASCII mode
  2. Binary mode: In "Binary mode", the sending machine sends each file bit for bit and as such the recipient stores the bitstream as it receives it.
In "ASCII mode", any form of data that is not plain text will be corrupted. When a file is sent using an ASCII-type transfer, the individual letters, numbers, and characters are sent using their ASCII character codes. The receiving machine saves these in a text file in the appropriate format (for example, a Unix machine saves it in a Unix format, a Windows machine saves it in a Windows format). Hence if an ASCII transfer is used it can be assumed plain text is sent, which is stored by the receiving computer in its own format. Translating between text formats might entail substituting the end of line and end of file characters used on the source platform with those on the destination platform, e.g. a Windows machine receiving a file from a Unix machine will replace the line feeds with carriage return-line feed pairs. It might also involve translating characters; for example, when transferring from an IBM mainframe to a system using ASCII, EBCDIC characters used on the mainframe will be translated to their ASCII equivalents, and when transferring from the system using ASCII to the mainframe, ASCII characters will be translated to their EBCDIC equivalents.

By default, most FTP clients use ASCII mode. Some clients try to determine the required transfer-mode by inspecting the file's name or contents, or by determining whether the server is running an operating system with the same text file format.

The FTP specifications also list the following transfer modes:

1. EBCDIC mode
2. Local mode

In practice, these additional transfer modes are rarely used. They are however still used by some legacy mainframe systems.

FTP and Web Browser
Most recent web browsers and file managers can connect to FTP servers, although they may lack the support for protocol extensions such as FTPS. This allows manipulation of remote files over FTP through an interface similar to that used for local files. This is done via an FTP URL, which takes the form ftp(s):// (e.g., ftp://ftp.gimp.org/). A password can optionally be given in the URL, e.g.: ftp(s)://:@:. Most web-browsers require the use of passive mode FTP, which not all FTP servers are capable of handling. Some browsers allow only the downloading of files, but offer no way to upload files to the server.

FTP and NAT Devices
The representation of the IPs and ports in the PORT command and PASV reply poses another challenge for NAT devices in handling FTP. The NAT device must alter these values, so that they contain the IP of the NAT-ed client, and a port chosen by the NAT device for the data connection. The new IP and port will probably differ in length in their decimal representation from the original IP and port. This means that altering the values on the control connection by the NAT device must be done carefully, changing the TCP Sequence and Acknowledgment fields for all subsequent packets.

For example: A client with an IP of 192.168.0.1, starting an active mode transfer on port 1025, will send the string "PORT 192,168,0,1,4,1". A NAT device masquerading this client with an IP of 192.168.15.5, with a chosen port of 2000 for the data connection, will need to replace the above string with "PORT 192,168,15,5,7,208".

The new string is 23 characters long, compared to 20 characters in the original packet. The Acknowledgment field by the server to this packet will need to be decreased by 3 bytes by the NAT device for the client to correctly understand that the PORT command has arrived to the server. If the NAT device is not capable of correcting the Sequence and Acknowledgement fields, it will not be possible to use active mode FTP. Passive mode FTP will work in this case, because the information about the IP and port for the data connection is sent by the server, which doesn't need to be NATed. If NAT is performed on the server by the NAT device, then the exact opposite will happen. Active mode will work, but passive mode will fail.

It should be noted that many NAT devices perform this protocol inspection and modify the PORT command without being explicitly told to do so by the user. This can lead to several problems. First of all, there is no guarantee that the used protocol really is FTP, or it might use some extension not understood by the NAT device. One example would be an SSL secured FTP connection. Due to the encryption, the NAT device will be unable to modify the address. As result, active mode transfers will fail only if encryption is used, much to the confusion of the user.

The proper way to solve this is to tell the client which IP address and ports to use for active mode. Furthermore, the NAT device has to be configured to forward the selected range of ports to the client's machine.

FTP over SSH
FTP over SSH refers to the practice of tunneling a normal FTP session over an SSH connection.

Because FTP uses multiple TCP connections (unusual for a TCP/IP protocol that is still in use), it is particularly difficult to tunnel over SSH. With many SSH clients, attempting to set up a tunnel for the control channel (the initial client-to-server connection on port 21) will protect only that channel; when data is transferred, the FTP software at either end will set up new TCP connections (data channels) which will bypass the SSH connection, and thus have no confidentiality, integrity protection, etc.

If the FTP client is configured to use passive mode and to connect to a SOCKS server interface that many SSH clients can present for tunneling, it is possible to run all the FTP channels over the SSH connection.

Otherwise, it is necessary for the SSH client software to have specific knowledge of the FTP protocol, and monitor and rewrite FTP control channel messages and autonomously open new forwardings for FTP data channels. Version 3 of SSH Communications Security's software suite, and the GPL licensed FONC are two software packages that support this mode.

FTP over SSH is sometimes referred to as secure FTP; this should not be confused with other methods of securing FTP, such as with SSL/TLS (FTPS). Other methods of transferring files using SSH that are not related to FTP include SFTP and SCP; in each of these, the entire conversation (credentials and data) is always protected by the SSH protocol.

FTP-like protocols
  • File Service Protocol (FSP)
  • FTPS (FTP/SSL), FTP run over SSL
  • Gopher, a hyperlinking anonymous FTP-like protocol
  • Secure copy (SCP), a protocol running over SSH
  • Simple File Transfer Protocol, the historic protocol RFC 913
  • SSH file transfer protocol (SFTP, SH-FTP, FTP/SSH), a protocol running over SSH
  • Trivial File Transfer Protocol (TFTP)
  • WebDAV, widely used extension of HTTP

Saturday, March 15, 2008

HTTP Explained

http://pwidodo.googlepages.com/httpexplained

How Internet Cookie work

http://www.howstuffworks.com/cookie.htm