Thursday, June 22, 2006

Ubuntu and Solaris 10 x86 on Laptop and Desktop

Cross pollination of OpenSolaris and GNU:
http://www.gnusolaris.org/gswiki

Tuesday, June 20, 2006

Compilation tunning tips

However, general tuning is easier than this, here's what I would suggest:

1. Run at -O to establish baseline performance
2. Run at -fast -xipo=2 -xtarget=generic[64]

If there's no difference (or no significant difference) in performance, then you can stop. [But still profile the application!]

If there is a difference, then I'd evaluate performance gains due to the following flags (some combinations may be missing):

3. -xO5
4. -xO5 -xalias_level=basic (for C) compatible (for C++)
5. -xO5 -xdepend
6. -xO5 -fsimple=2 -fns -xlibmil -lmopt
7. -xO5 -xipo=2

I think that covers the bulk of the things that get enabled at -fast. Hopefully from these runs you'd be able to isolate a set of flags which gives you performance.

You might also want to look into profile feedback for codes which contain lots of branch instructions (or calls).

Obviously profiling the application (eg perhaps with spot http://cooltools.sunsource.net/spot/) will give you insights into what the actual performance issues are, and these insights can guide you to selecting appropriate compiler flags.

compiler -fast and optimization on different systems

If Sun studio taking top down or bottom up approach for backend optimization ? One more thing to share I am considering if -fast macro expansion is NP hard. In addition, should we further approximate the computation procedures without generating new sub-problem running on different underline system. Or we should consider different algorithm to conquer the problem.

A simplification on Compiler process

1. Parsing ("front end"). Breaks down the code to elementary operations, and generates debug information. The C++ front end (but not the C front end) inlines some functions that were explicitly or implicitly declared inline. Both C and C++ can mark functions for the back end as "please generate inline."

2. Code generation ("back end"). Generates code, with variable levels of optimizing. Function inlining can occur at -xO3 by request, and at -xO4 or -xO5 even if not requested.

The -g option affects primarily the front ends. The back ends disable some optimizations when -g is in effect. The -g option disables front-end inlining by the C++ compiler; the -g0 option enables front-end inlining.

Most of the remaining optimization options primarily affect the back end, and are the same for C and C++. A few options are available only in C. You need to check the documentation for the compiler version you use, because new options are added from time to time.

The -fast option is really a macro that expands to a series of options based on the details of the system running the compiler. If you run the resulting program on a different machine, results will be sub-optimal, maybe even slower than if it were not compiled with -fast.

The -xOn options select optimizations that are useful across a range of systems.

Boot NFS from NG-Zone

Booting a zone over NFS (namely where its root file system is on the NAS device) is not supported at the current time. There are um...interesting workarounds like using lofi(7D) available.Boot support over NFS is definitely something we want to support in the future, Just to make it clear though - a Solaris 10 system acting as a NAS device can itself host its own non-global zones. They just need to boot off of local file systems on the system.

A classic use case for Zone booting over NFS could be identified within grid computing environment
in order to achieve resource allocation.It is critical since I consider zone is considered as
resource fabric to access large (TB) over underline network NFS. Otherwise grid service providers need to implement quite Transactional grid enabled network file system management services. Of course, it has to be zone aware.

Friday, June 16, 2006

Middleware enters ESB age and SOA is ready

All Middleware vendors adopt ESB solution.
SOA age is ready. ESB next standard in
the Java EE stack ?

Thursday, June 15, 2006

A real world problem and algorithmic analysis

Problem:
An engineer using Toshiba Tecra M2 laptop, and it is failing on him.The problem seems like either power adaptor or the battery pack, but he needs to determine which before he tried to purchase any replacement. For that, He would like to do simple test with existing ones.

Objective:

fix the problem instead of bringing up to run.

Algorithm Analysis and Design:


0x00000001 --- AC Adapter
0x00000010 --- Laptop Main

Tested Unit:
Parts[0] <---- 0x00000001
Parts[1] <---- 0x00000010

Good Laptop as Instrumentation Tool:
Tool[0] <-----0x00000011
Tool[1] <-----0x00000100


Fix-Laptop(Parts, Tool) return 0-1
DefectParts <---- Test-Part (Parts, Tool)
Go to Dealer for maintainance
done <---- 1
return done


Test-Part (Parts, Tool) return DefectParts
Success = Connect(Parts[1], Tool[0]);
if Success
then DefectParts[0] <---- Parts[0]
else do Success = Connect(Parts[1], Tool[1]);
then DefectParts[0] <---- Parts[1]
else do inspection again ensure no contact issue
else DefectParts <----- Parts
return DefectParts

Solution:

It seems unicast to local vendor would be optimal
algorithm than multicast over the smtp overlay
network in term of complexity, cost and completeness.
Since We need to go to dealer anyway, why do
test there and get fixed part right away ?

CDROM access from NG-Zone

How to access cdrom drive from the zone

(1) Try the followings in GZ

/etc/init.d/volmgt start

Make sure in LZ the following SMF services are online too:

svcadm enable svc:/network/rpc/bind:default
svcadm enable svc:/network/rpc/smserver:default

Check the following in case the previous commands don't succeed
Run prtconf to see if the "sd" driver is attached to the cdrom or not.
If not, rem_drv sd and add_drv sd to see if sd can be attched to the cdrom.
If attach fails, then you have a problem. No matter what you are trying, your cdrom cannot be mounted at all.

The bottom line, the "sd" driver attached to the cdrom hardware.


(2) To add a CD-ROM:

run zonecfg and add the foolowing statements:

add fs
set dir=/cdrom
set special=/cdrom
set type=lofs
set options=[nodevices]
end

report system configuration with mdb

Run as root

echo "::prtconf" | mdb -k


It will report all device configuration

Wednesday, June 14, 2006

T1 and e1000g

Enabling e1000g
==============

The Ontario motherboard has Intel Ophir chip that can be used with ipge or e1000g network drivers.
Typically the factory default driver is ipge. If you want to exercise the e1000g driverinstead of
the ipge, please follow the following steps.

How to switch to e1000g driver from factory default ipge driver:
========================================================
1) Edit /etc/rc2.d/S99bench. Plumb e1000g and comment out the plumbing of ipge.
using the command ifconfig e1000g plumb

2) In a separate window on the same machine, run, prtconf -pv and see what
compatible vendor ids are shown for the network interface
You will see entries such as :
>>
compatible: 'pciex8086,105e.108e.105e.6' + 'pciex8086,105e.108e.105e' +
'pciex8086,105e.6' + 'pciex8086,105e' + 'pciexclass,020000' +
'pciexclass,0200'
<<

Check the driver aliases file for ipge entries. You will see values such as :
ipge "pciex8086,105e"
ipge "pciex8086,105f"
ipge "pci8086,105e"
ipge "pci8086,105f"

Check if any of these vendor ids match with vendor-product ids already listed
for e1000g driver. Backup /etc/driver_aliases file to /etc/driver_aliases.ipge Replace all ipge to e1000g in /etc/driver_aliases

3)
Backup original /etc/path_to_inst file to path_to_inst.ipge.
Now, change all the ipge entries in path_to_inst file to e1000g.
Note: the port numbers change too.

port 1 e1000g --> port 0 ipge
port 3 e1000g --> port 1 ipge
port 0 e1000g --> port 2 ipge
port 2 e1000g --> port 3 ipge

Check out the diff below and edit your path_to_inst accordingly.

testmachine> diff path_to_inst path_to_inst.ipge
10,11c10,11
< "/pci@780/pci@0/pci@1/network@0" 1 "e1000g"
< "/pci@780/pci@0/pci@1/network@0,1" 3 "e1000g"

> "/pci@780/pci@0/pci@1/network@0" 0 "ipge"

> "/pci@780/pci@0/pci@1/network@0,1" 1 "ipge"

18,19c18,19
< "/pci@7c0/pci@0/pci@1/network@0" 0 "e1000g"
< "/pci@7c0/pci@0/pci@1/network@0,1" 2 "e1000g"

> "/pci@7c0/pci@0/pci@1/network@0" 2 "ipge"

> "/pci@7c0/pci@0/pci@1/network@0,1" 3 "ipge"

4) Copy /etc/hostname.ipge2 to /etc/hostname.ipge2.bak
Rename /etc/hostname.ipge2 to hostname.e1000g0

5) Reboot the machine

6) When the machine comes up now, run ifconfig -a. You should be able to see
e1000g0
7) Check the inet and netmask for e1000g0 and correct it if necessary

8) Plumb other ports and set inet and netmasks for them also using:
#ifconfig e1000g0 inet netmask up
9) Setup default gateway (Get default gateway using netstat -nr)
#route add default

10) Check cables and make sure leds are green
You should now be able to ping through e1000g on all the interfaces


How to use a new e1000g driver on factory installed Ontario:
=============================================================

Obtain the latest e1000g driver files: e1000g and e1000g.conf
(You may want to contact the e1000g driver team)


1) Copy driver and conf file.
copy e1000g binary to /kernel/drv/sparcv9/
copy e1000g.conf file to /kernel/drv/

2) Backup /etc/driver_aliases to /etc/driver_aliases.ipge

3) Modify /etc/driver_alias

4) Replace all the ipge to e1000g.

5) Backup /etc/path_to_inst to /etc/path_to_inst.ipge

6) Modify /etc/path_to_inst by replacing ipge with e1000g.
Note, the port numbers will change too.
Port for ipge1 becomes e1000g3 and ipge2 becomes e1000g0.

7) Modify /etc/name_to_major. Add a line at end "e1000g 267".
(Go to the last line and select the number that is consecutively higher).

8) Run: touch /reconfigure

9) cp /etc/hostname.ipge2 /etc/hostname.e1000g0
10) Edit /etc/rc2.d/S99bench to plumb e1000g and comment out ipge.
11) Reboot machine.

12) When the machine comes up now, run ifconfig -a. You should be able to see
e1000g entry

13) Check the inet and netmask for e1000g0 and correct it if necessary

14) Plumb other ports and set inet and netmasks for them also using:
#ifconfig e1000g0 inet netmask up

15) Setup default gateway (Get default gateway using netstat -nr)
#route add default

16) Check cables and make sure leds are green
You should now be able to ping through e1000g on all the interfaces

T1 CPI

(7) The instruction execution resulting in overlapping latency which leads to the memory model of T1 addresses the effectiveness contributed with or without memory stalls across the memory
hierarchy. Empirical data set indicates the problem size and requires further investigation on the hidden factors contributing the CPU efficiency.

(6) Someone may agree. A CPI of >=4 as tested on a Niagara would indicate linear thread scalaing, but the same data found on an USIII would not necessarily lead to the same conclusion. 1 - 2 CPI on an USIII could actually be >=4 CPI on a T1 because of the USIII's superscalarness. The amount of thread level paralellism in an instruction stream is somewhat limited. So the real question is, is it possible for a *realistic* instruction stream to have a number of stalling instructions that would give >4 CPI on a T1 but closer to 1 CPI on an USIII, due to the USIII's superscalarness masking the stalls? I'm not so convinced that this is true. But data would be good.

(5) Someone questioned why a cpi > 4[as seen on a USIII] is required for a workload to scale linearly on T1."most of the kernels don't meet the T1 requirement of a cpi of 4 to get thread scaling".USIII has instruction level parallelism so (theoretically) a CPI greater than 4(as seen on USIII) should not be a necessary condition to linearly scale on T1.


(4) Data on which specint tests contain heavy FP. I don't have the data, but I suspect the twolf test also has decent FP as it's another place and route test like vpr.Eon is a graphics visualization test. Also see Brian's comments about Niagara's CPI for specint, indicating that even if you discount the performance on the FP heavy workloads, specint still won't do as well as a "real world" workload that has average cpi > 4.

(3) the SPECint_rate FP data set

Percent fp...
vpr dataset 1 -> 5.6%
dataset 2 -> 8%
eon dataset 1 -> 15.9%
dataset 2 -> 15.2%
dataset 3 -> 16.3%
Only 0.1% of the instructions in eon are sqrt, so fixing
sqrt will help single core, but not significantly change
the rate result.

CPI
gzip ds1 -> 1.14
ds2 -> 1.10
ds3 -> 0.97
ds4 -> 0.97
ds5 -> 1.17
vpr ds1 -> 1.34
ds2 -> 2.37
gcc ds1 -> 2.19
ds2 -> 1.37
ds3 -> 1.50
ds4 -> 1.64
ds5 -> 1.46
mcf ds1 -> 5.81
crafty ds1 -> 1.00
eon ds1 -> 1.20
ds2 -> 1.25
ds3 -> 1.30
perlbmk ds1 -> 1.27
ds2 -> 1.08
ds3 -> 1.74
ds4 -> 1.03
ds5 -> 1.07
ds6 -> 1.04
ds7 -> 1.06
gap ds1 -> 1.46
vortex ds1 -> 1.38
ds2 -> 1.24
ds3 -> 1.39
bzip2 ds1 -> 1.11
ds2 -> 0.91
ds3 -> 0.97
twolf ds1 -> 1.94

(data collected by Darryl Gove on a US3 1056MHz system)

So, for this benchmark, most of the kernels don't meet the T1
requirement of a cpi of 4 to get thread scaling. That, along
with a single issue processor make it impossible to get good
numbers on this benchmark.

So, the problem really is, int_rate doesn't stall on memory enough
for the T1 processor.

I noticed a reply from you to niagara-interest saying that according to folks at SAE, specint is ~20% floating point. Do you know where I might be able to find a breakdown of FP % for each of the 12 benchmarks.

I have a partner that uses specint results to compare platforms internally and is doing some Niagara testing. They're aware that specint contains floating point instructions and are willing to take suggestions from us on how the different benchmarks should be weighted to emphasize integer performance (I'm hoping, of course, that there are some benchmarks with little to no FP).

D consumer and libtrace api

How to use libdtrace api's to interact with the dtrace
subsystem. I just want to use certain methods from within
a 'c' program.


With the user land system calls, which D consumer would be in
your mind for call back invokation from your probes ?

Thursday, May 04, 2006

Probabilistic analysis and Random algorithm

(1) Cost Model and Running Time Model are different
(2) The Technique which is used to analyze Cost Model and Running Time Model is the same.
(3) It is to count the number of operation execution in the routine algorithm
(4) Probabilistic analysis is one of a technique which fits in the use case where there is a input distribution. However, if we can not describe a reasonable input distribution, we can not use probabilistic analysis.
(5) In many case, we only can know a little bit about the distribution of the input but can not model the knowledge of the input distribution.
(6) Random algorithm is controlled not only by input distribution but also random-number generator
so that we can have a level of control which I do not need to guess and make assumption that input
comes in with random order instead we choose input randomly to ensure input is definitely random.
We used to call a pseudorandom number generator --- deterministic algorithm returning random number
(7) Probabilistic analysis impose a distribution rather than assuming a distribution of inputs to
the development of the randomized algorithm
(8)

Friday, April 07, 2006

Fidelity

As the example in the previous section illustrates, the data accessed
by an Odyssey application may be stored in one or more gcncralpurpose
repositories such as file servers, SQL servers, or Web
servers. Alternatively, it may be stored in more specialized rcposltories
such as video libraries, query-by-image-content databases, or
back ends of geographical information systems.
The constraints of mobility complicate data access from such
servers. Ideally, a data item available on a mobile client should be
indistinguishable from that available to the accessing application if
it were to be executed on the server storing that item. But this corrcspondencemay
be difficult to preserve as resources become scarce;

Thursday, March 30, 2006

disable zlogin

Currently have scripts
running during the provisioning of zones to populate
SOE packages and preliminary configuration that need to
be done prior to granting access to anyone, including
any global administrator until after the configuration
is complete.

To avoid zlogin to local zone,
Add the following lines to the
zone's /etc/pam.conf just about
the "other auth" lines

#
# disable zlogin
zlogin auth required pam_deny.so.1

Tuesday, March 28, 2006

Solaris motd issue

(1) Here is how login(1) routine works
/usr/sbin/quota (check quota)
/bin/cat -s /etc/motd (print motd)
/bin/mail -E (check mail)

(2) Here is how mibiisa(1M) works as SNMP agent utility
motd is part of sunsystem group for general
system information reporting. The first line
of /etc/motd. (string[255])

(3) For JES link, I have JESQ4 on my system, it does
not show the link. Have you check which component
create the link ?

(4) A few line D code below may help you to discover the issue


performance analysis for the alogrithm
counts the cost the steps of the random
access machine which is to modeled for
the instrumentation.

Consequently, higher level syscall
instrumentation
symlink(*char* *target , *char* *linkname)
tracing seems friendly for the implementation
of the code. It does not give a plus to performance
Therefore, I would keep the routine as close to
the I/O layer in order to mini the cost of the
delegation of the layered kernel architecture.

I would suggest to add directive as condition
rule to point to the path to the motd in order
to filter out the I/O.

symlink(2)does
the link and rename only. AI and Algoritm
calculation does make sense. The
implementation of my AI and Algorithm
is enhanced as code below.

Please let me know if it works on your system



#! /usr/sbin/dtrace -s
#pragma D option quiet

dtrace:::BEGIN
{
printf("%15s %40s\n", "Executable", "LinkFileName");
}
/* Please note input here is link file name not path */
fbt::fop_symlink:exit
/stringof(args[1]) == $$1/
{
printf("%15s %40s\n", execname,stringof(args[1]));
}

In addition, you can seperate the R/W to further
narrow down the report.

Here is a script that will print the time, name of the executable,
and ptree output when anyone tries to link /etc/motd

#!/usr/sbin/dtrace -wqs

syscall::symlink:entry
/basename(copyinstr(arg1))=="motd"/
{
printf("Caught the culprit\n");
printf("%20s\t %-20Y\n", "Time",walltimestamp);
printf("%20s\t %-10d\n", "Process id",pid);
printf("%20s\t %-20s\n", "Name of Executable" ,execname);
stop();
system("ptree %d",pid);
system("prun %d",pid);
}

Also if they want to use DTrace to automatically avoid the process
from creating the link they can use the script below. This would cause
any link to /etc/motd to become a link to /tmp/motd and then remove the
/tmp/motd file.

#!/usr/sbin/dtrace -wqs

syscall::symlink:entry
/copyinstr(arg1)=="/etc/motd"/
{
printf("Caught the culprit\n");
printf("%20s\t %-20Y\n", "Time",walltimestamp);
printf("%20s\t %-10d\n", "Process id",pid);
printf("%20s\t %-20s\n", "Name of Executable" ,execname);
copyoutstr("/tmp/motd",arg1,9);
stop();
system("ptree %d",pid);
system("prun %d",pid);
system("rm /tmp/motd");
}

Monday, March 27, 2006

Second hit on OS kernel architecture model design

As computer science illustrates:
a goal based model agent does the routine
perception, rule matching and goal mapping
in order to approximate the actions to deal
with the subset of the unobserved conditions.
Consequently, what will be the input of the
change ? kernel, specifically core kernel
modules seems relative stable. What else ?
security patches ? certified third party
drivers ? system library ? user lander
applications ? Is this the time to review the
challenge traditional open system kernel
architecture model design to deal with
complexity ? It is the common control,
plan and game algorithms, models along with
economic models inspires CS scientists
to review the challenges for OS vendors.
Microsoft is not alone. I will be valuable
to investigate the asymptotical notation
for both best cases and worst cases.

On the another hand, what will be business
intelligence to protect OS vendor market
shares along with the user land applications
for OS players ? Does the user land application
continue lock end users ? What will be the
economic and innovative delivery model for
both OS vendors and application providers ?
What OS vendors can really get from open
sources or open services environment ?

Many and Many questions and thoughts ?!


http://www.nytimes.com/2006/03/27/technology/27soft.html?hp&ex=1143522000&en=1c725e1c50ae8d6c&ei=5094&partner=homepage

Tuesday, March 21, 2006

System Event Handling for Non Global Zone via GPEC event queue channel

System event is not allowed in NG-zone. However, system calls
via sysevent 3SYSEVENT is the working solution

sysevent_bind_handle(3SYSEVENT) – bind or unbind subscriber handle
sysevent_free(3SYSEVENT) – free memory for sysevent handle
sysevent_get_attr_list(3SYSEVENT) – get attribute list pointer
sysevent_get_class_name(3SYSEVENT) – get class name, subclass name, ID or buffer size of event
sysevent_get_pid(3SYSEVENT) – get vendor name, publisher name or processor ID of event
sysevent_get_pub_name(3SYSEVENT) – get vendor name, publisher name or processor ID of event
sysevent_get_seq(3SYSEVENT) – get class name, subclass name, ID or buffer size of event
sysevent_get_size(3SYSEVENT) – get class name, subclass name, ID or buffer size of event
sysevent_get_subclass_name(3SYSEVENT) – get class name, subclass name, ID or buffer size of event
sysevent_get_time(3SYSEVENT) – get class name, subclass name, ID or buffer size of event
sysevent_get_vendor_name(3SYSEVENT) – get vendor name, publisher name or processor ID of event
sysevent_post_event(3SYSEVENT) – post system event for applications
sysevent_subscribe_event(3SYSEVENT) – register or unregister interest in event receipt
sysevent_unbind_handle(3SYSEVENT) – bind or unbind subscriber handle
sysevent_unsubscribe_event(3SYSEVENT) – register or unregister interest in event receipt

How to resolve fsflush overhead as large sized memory mapped

To reduce the I/O on Solaris platform, Solaris does
offer virtual file system which is memory-based file systems
that provide access to kernel specific resources.
As the name indicates,virtual file systems do not
use file system disk space. However, tmpfs use the
swap space on a disk. tmpfs is the default file system
type for the /tmp directory in the Solaris.

Since uses local memory for file system reads and writes,
it has much more lower latency than using classic Solaris
UFS. I/O performance can be enhanced by reduced I/O
to a local disk or across the network in order to significantly
speed up their creation, manipulation etc. Therefore, tmpfs
can be utlized for the memory mapping.

Files in TMPFS file systems are votial. The files will be dispeared
as the file system is unmounted and when the system is shut down
or rebooted. Files can be moved into or out of the /tmp directory.
This means KTS needs to ensure completed process image backup
as normal fsflush does.

Please note that tmpfs uses swap space for pageout. Process will
be executed as system does not have enough swap space. Which means
it requires larger swap space for pagout.


Other than Solaris built-in kernel modules, Sun Storage Cache also can
be utlized to reduce the latency of the I/O activities.


To ensure no paging on Solaris platform, shmop(2) with shmget(2)
and Intimate Shared Memory variant of System V shared
memory. ISM* *mappings are created with the SHM_SHARE_MMU flag.
This locks down the memory used. Then just read the file
into shared memory. But this may result in code change.



If this is not an option you can tune the flusher with system parameters

set segspt_minfree
set swapfs_minfree
set lotsfree
set desfree
set minfree

You can also postpone the time between cleaning with

set autoup

Monday, March 20, 2006

Solaris Page-Demand Memory Management

Pertaining Solaris Kernel Proc mgt and Memory Mgt
architecture, Heap segment within Process Virtual
Address Space is allocated for user land data structure
righ above executable data segment of the user land
DB process and grow with the libc.so.1 system
library call such as malloc(3c) which malloc_unlocked does
the dirty work to allocate holding blocks or ordinary
blocks for user land process. If there is no block sbrk(3c) is
called.

It is transparent zero-fill-on-demand memory page
allocation because of the page fault.
Page memory is allocated for the process heap and
space becomes a permanently allocated block.

(1) However, the allocation will not shrink until
process exits.
(2) Page scanner daemon runs to page out memory page
per LRU due to the shortage of memory

This is the core of on demand page memory management
architecture of Solaris Operating System.

As for the free(3c) which free_unlocked is doing the dirty
work to mark the address space as free list
for later use but not to return address space to memory
resource managed pool.

Solaris Basic Library

These default memory allocation routines are safe for use
in multithreaded applications but are not scalable.
Concurrent accesses by multiple threads are single-threaded
through the use of a single lock. Multithreaded applications
that make heavy use of dynamic memory allocation should be
linked with allocation libraries designed for concurrent access,
such as libumem(3LIB) or libmtmalloc(3LIB). Applications that
want to avoid using heap allocations (with brk(2)) can do so
by using either libumem or libmapmalloc(3LIB). The allocation
libraries libmalloc(3LIB) and libbsdmalloc(3LIB) are
available for special needs.